File size: 9,229 Bytes
5d08972 28ed187 5d08972 28ed187 5d08972 28ed187 5d08972 28ed187 5d08972 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 | # pose6daug β augmentation, dataset build, training and evaluation scripts
The scripts behind the MimicGen and VACE augmentation baselines for RoboCasa
PickPlaceCounterToCabinet, and the GR00T 1.5 fine-tuning and exact-replay evaluation run on
top of them. Archived as-run: paths are absolute and point at one particular machine, so treat
this as a record of the procedure rather than a turnkey package. See **Paths to change** below.
```
augment/mimicgen/ generate episodes with MimicGen, render preview videos
augment/vace/ VACE object-swap augmentation (assignment builder, runner, GT masks)
augment/actaug/ actaug episode folders -> gr00t_views (no simulator needed)
dataset/ convert generated episodes into a gr00t_views (LeRobot v2.1) dataset
train/ GR00T 1.5 fine-tuning launcher
eval/ exact-state replay evaluation -- read eval/EVALUATION.md for the protocol
ops/ checkpoint retention, optimizer pruning, eval-on-checkpoint triggers
tools/ dataset sanity checks
```
## Pipeline
```
MimicGen generation βββΊ demo.hdf5 (MuJoCo states + actions, no pixels)
β
βββββββββββ΄βββββββββββ
βΌ βΌ
preview mp4 (3 views) gr00t_views dataset βββΊ fine-tune βββΊ exact-replay eval
```
VACE skips the simulator: it repaints pixels and keeps the source trajectory, so its builder
copies the source parquet instead of replaying.
### 1. Augment
`augment/mimicgen/run_generation.sh` launches one MimicGen process per worker from a JSON
config per (object, worker); `config_template.json` is the template. Two knobs matter:
- `guarantee=false` makes `num_trials` mean **attempts**, not successes. Giving every object
the same attempt budget is what produces the generator's natural yield instead of a
per-object quota.
- `obj_registries` must include `aigen` for objects under `aigen_objs/` (e.g. `wine_5`), or
`sample_kitchen_object_helper` raises a bare `ValueError`.
`snapshot_episode_times.sh` records the per-episode temporary filenames while they exist:
MimicGen writes each success to `tmp/date_..._time_HH_MM_SS.hdf5`, then `merge_all_hdf5`
sorts by timestamp and deletes the folder. That snapshot is the only record of when each
episode was produced, and `select_by_generation_time.py` joins it back to order episodes
globally across workers. The poller can only ever miss a worker's **last** file, which the
selector pads with the merged file's mtime.
### 2. Build the dataset
`dataset/render_to_gr00t.py` replays each episode in MuJoCo, renders three cameras and writes
parquet + videos + meta. `dataset/run_convert.sh` shards it (each shard needs its own
`GEN_DIR` and `DATASET_OUT`, or the glob picks up the others), then
`merge_gr00t_view_datasets.py` merges and `repair_mimicgen_task_ids.py` fixes task ids.
`repair_mimicgen_task_ids.py` is not optional: the writer stores every parquet task column as
0 while `episodes.jsonl` holds the intended language, so without it every episode trains as
task 0 and the language conditioning silently collapses.
### 3. Train
```bash
DATASET_PATH=<gr00t_views dataset> DATASETNAME=<name> \
GPUS=2,3 PER_GPU_BATCH=32 MAX_STEPS=30000 SAVE_STEPS=5000 \
bash train/train_groot15_single_dataset.sh
```
Fine-tunes the action-head projector and diffusion head from a base checkpoint; the backbone
stays frozen. `RESUME=1` picks up the newest checkpoint in the output directory.
### 4. Evaluate
The full protocol -- episode set, seeds, which checkpoints to compare, how to read the
stage flags -- is in [`eval/EVALUATION.md`](eval/EVALUATION.md). The short version:
```bash
MODEL_PATH=<checkpoint> MYGROOT_ROOT=<myGR00T tree> \
EVAL_CLIENT=<eval/eval_robocasa_replay_state_grasp.py> \
REPLAY_STATE_ROOT=<replay set> N_EPISODES=160 GPUS=4,5,6,7 POLICY_SEED=12345 SEED_BASE=42 \
bash eval/eval_groot15_exact_replay.sh
```
Each episode restores a saved scene XML and flattened MuJoCo state, so runs are comparable
across checkpoints. Two independent seeds:
- `POLICY_SEED` β the policy server's global RNG **and** the per-step action seed
(`policy_seed + episode_index Γ stride + step`). This is the one that changes the rollout.
- `SEED_BASE` β the env seed, `SEED_BASE + worker_id` per worker. Exact replay overwrites the
scene immediately after, so it should not affect the initial state.
Worker count follows the GPU list, and episodes are split evenly across workers β so the same
`SEED_BASE` with a different GPU count gives each episode a different worker seed.
Client variants:
| file | adds |
|---|---|
| `eval_robocasa_replay_state_3view.py` | saves the wrist (ego) view; `RecordVideo` only captures `robot0_agentview_center` |
| `eval_robocasa_replay_state_grasp.py` | the above, plus `--episode_indices` for an arbitrary subset, and per-episode `grasped` / `lifted` / `in_cab` stage flags |
Stock RoboCasa success is `obj_inside_of(cab) and gripper_obj_far` β a single boolean, which
says nothing about where a failed episode broke down. The stage flags come from
`_check_grasp`, a 3 cm rise in the object's body z, and `OU.obj_inside_of`.
`organize_videos.py` regroups the output into one folder per global episode
(`rollouts/episode_NNNNNN/{center.mp4, wrist.mp4, info.txt}`); `RecordVideo` names files by
worker-local index, which does not match the global index.
## The action-order bug β check any dataset before training on it
RoboCasa's simulator exports actions **arm-first**:
```
[eef_pos(3), eef_rot(3), gripper, base(3), torso, base_mode]
```
A gr00t_views dataset declares them **base-first**:
```
base_motion[0:4] control_mode[4:5] eef_pos[5:8] eef_rot[8:11] gripper_close[11:12]
```
A builder that re-derives actions from a simulator rollout has to reorder; one that copies
rows from an existing LeRobot dataset does not. Copying them through unchanged puts arm motion
on the base-motion channel and `base_mode` on the gripper. Nothing errors, the loss converges
to a small value, and the policy drives the base away from the counter and never closes the
gripper β **0/160** on an exact-replay eval whose base checkpoint scored 11/160.
```python
actions = raw_actions[:, [7, 8, 9, 10, 11, 0, 1, 2, 3, 4, 5, 6]]
```
`tools/check_action_layout.py` tells the two layouts apart from the data alone (the mobile
base never moves in this task, so the constant dimensions give it away), and
`dataset/repair_action_order.py` fixes an already-built dataset in place β parquet plus the
action entry of `meta/stats.json`, no re-render. It refuses to run on a dataset that does not
look like simulator order, so it cannot be applied twice. Read `tools/ACTION_LAYOUT.md` first.
## Paths to change
Every script hard-codes absolute paths from the machine this was run on. At minimum:
| what | appears as |
|---|---|
| RoboCasa / robosuite checkouts | `/lp-dev/jonghoon/robocasa_calib/repos/...` |
| MimicGen env + augmentation code | `/lp-dev/jonghoon/mimicgen_augment/...` |
| myGR00T tree and conda envs | `/data/minha/pose6daug/train_robocasa/myGR00T`, `/data/nvidia/gripper_augmentator/conda-envs/...` |
| base checkpoint | `/lp-dev/jonghoon/myGR00T_outputs/pnpcountertocab_all502_.../checkpoint-60000` |
| replay sets and eval output | `/lp-dev/jonghoon/isaac-gr00t/eval_results/...` |
| generation scratch | `/tmp/claude-.../scratchpad/...` |
`eval/eval_groot15_exact_replay.sh` defaults `MYGROOT_ROOT` to a path that no longer exists;
pass it explicitly. No credentials are embedded β the training launcher reads
`WANDB_API_KEY` from an env var or a file path you supply.
## Attribution
`eval/eval_groot15_exact_replay.sh` and the upstream of `dataset/render_to_gr00t.py` and the
eval clients come from the pose6daug project's shared tree; the copies here carry the fixes
described above (action reorder, mesh-swap bbox fix, target sharding, wrist-view capture,
stage logging).
## actaug
`augment/actaug/build_actaug_gr00t.py` converts actaug's per-episode folders
(`actions.npy`, `states.npz`, `left/right/wrist.mp4`, `meta.json`) into a gr00t_views dataset.
Like the VACE builder it needs no simulator -- the videos and states already exist.
Two things it handles:
- **Instruction rewrite.** actaug keeps the *source* episode's language (65 distinct strings,
e.g. "Pick the wine ..." on a SoapDispenser010 episode). The augmented object is the one in
the folder name, so the instruction is regenerated from that, giving the same 7 task strings
the other datasets use.
- **Action order.** actaug already writes base-first, matching `modality.json`, so no reorder
is needed -- unlike the MimicGen path. Its actions do exercise the mobile base in some
episodes, which the other datasets never do, so `tools/check_action_layout.py` cannot
classify it (that check assumes a static base). Verify by hand there instead.
## Related artifacts
- `mlnha/mimicgen-batch64-30k-ckpts` β checkpoints from the fixed-action-order run
- `mlnha/vace-batch64-30k-ckpts` β VACE run checkpoints
- `mlnha/mimicgen-pi05-aug256` β MimicGen augmentation for the pi0.5 hard-object set
|