File size: 5,497 Bytes
28ed187
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
# Evaluation protocol

Exact-state replay on RoboCasa PickPlaceCounterToCabinet. Every run restores the same 160
saved scenes, so checkpoints are directly comparable.

## The episode set

`pnpcountertocab_mimicgen8_exact160_replay_policyseed12345_20260506` — 160 episodes,
**20 per object** for the 8 augmentation targets:

```
  0~ 19  donut/donut_5              80~ 99  steak/steak_8
 20~ 39  jar/Jar023                100~119  syrup_bottle/SyrupBottle005
 40~ 59  measuring_cup/MeasuringCup009  120~139  teapot/teapot_6
 60~ 79  soap_dispenser/SoapDispenser010  140~159  teapot/teapot_7
```

Drawn from a 1,160-episode base-60k rollout collection over 58 objects, keeping the episodes
for these 8. That collection scored 45.0% overall (466/1035 completed); on these 160 it scored
15/160, so this is a deliberately hard slice — a policy doing well elsewhere still looks low
here. Each episode ships `model_xml_gz`, `state_npz` and an `ep_meta` pickle; the wrapper
restores the XML and the flattened MuJoCo state, so the initial condition is identical across
runs rather than re-sampled.

## Settings — hold these fixed

```bash
MODEL_PATH=<checkpoint>                     # the only thing that should vary
MYGROOT_ROOT=<myGR00T tree>                 # the default path in the launcher is dead
EVAL_CLIENT=<.../eval_robocasa_replay_state_grasp.py>
REPLAY_STATE_ROOT=<.../pnpcountertocab_mimicgen8_exact160_replay_policyseed12345_20260506>
N_EPISODES=160 EXPECTED_REPLAY_STATES=160
POLICY_SEED=12345 SEED_BASE=42 GPUS=<four entries>
```

`POLICY_SEED` is the one that changes rollouts: it seeds the policy server's global RNG and
the per-step action seed (`policy_seed + episode_index x stride + step`). It matters — on one
checkpoint, 54321 gave 2/160 and 12345 gave 6/160, a 3x swing from the seed alone. **Do not
read small differences between single-seed runs as a result.**

`SEED_BASE` seeds the env (`SEED_BASE + worker_id`), but exact replay overwrites the scene
immediately after reset, so it should not affect the initial state — unverified, so keep it
at 42 anyway.

Worker count follows the GPU list and episodes split evenly across workers, so the GPU count
changes which worker seed each episode gets. Keep **four workers**. To put four workers on two
cards, repeat them: `GPUS=2,3,2,3`. Each server needs roughly 20 GB, so two servers per card
needs ~42 GB free — check before launching if training shares those cards, or the OOM takes
the training down too.

## Which checkpoints

Compare like with like. The runs completed so far:

| model | checkpoint | steps | note |
|---|---|---|---|
| base | `pnpcountertocab_all502_gbs64_wandb_60k_save5k_.../checkpoint-60000` | 60k | the checkpoint every augmentation run fine-tunes from |
| MimicGen | `mimicgen_natural_256_fixed/checkpoint-30000` | 30k | ~28 epochs over 256 episodes |
| MimicGen (short) | `mimicgen_natural_256_short/checkpoint-1500` | 1.5k | same data and config, ~1.4 epochs |
| VACE | `vace_batch64_30k_from_hf/checkpoint-30000` | 30k | |
| actaug | `actaug_256/checkpoint-1500` | 1.5k | |

Always evaluate the base checkpoint under the same harness rather than quoting a number from
the run that produced the scenes. Those two are not the same measurement: the collection run
recorded 15/160 while exact replay of the same episodes with the same policy seed gives
**11/160**, so 6.9% is the baseline to compare against, not 9.4%.

Step counts differ between runs, so a 1.5k checkpoint against a 30k one confounds training
length with the data. Evaluate the same step count on both sides before attributing a
difference to the dataset.

## Reading the output

```
<EVAL_OUTPUT_BASE>/<EVAL_RUN_NAME>/
  summary.json              n_episodes, n_success, success_rate, and the full run config
  simulation_results.csv    per episode: success, reward, policy_steps, video_path,
                            layout_id, style_id, language, target_mesh_keys
  episodes.jsonl            merged per-episode rows (the stage flags live here)
  workers/worker_NN_seed_S/ per-worker logs, videos, ep_meta pickles
  rollouts/episode_NNNNNN/  center.mp4 + wrist.mp4 + info.txt, after organize_videos.py
```

`policy_steps x action_horizon` equals `sim_step_budget` (47 x 16 = 752) on any episode that
ran out the clock, which is what failure looks like — there is no early abort.

### Stage flags

Stock RoboCasa success is `obj_inside_of(cab) and gripper_obj_far`: one boolean, which cannot
say where a failure broke down. `eval_robocasa_replay_state_grasp.py` adds three per-episode
flags, each latched if it ever held during the rollout:

| flag | source |
|---|---|
| `grasped` | `env._check_grasp(gripper, object_geoms)` |
| `lifted` | object body z rose more than 3 cm above its start height |
| `in_cab` | `OU.obj_inside_of(env, "obj", env.cab)` |

They localise the failure. On one run: 45 grasped, 29 lifted, 27 in cabinet, 25 succeeded —
a third of successful grasps were dropped during the lift, while nearly everything lifted went
on to finish. Without the flags all of that reads as a single success rate.

## Sampling a subset

`--episode_indices` takes an arbitrary list, for a quick read before committing an hour per
checkpoint. Five per object:

```bash
python3 -c "print(','.join(str(b+i) for b in range(0,160,20) for i in range(5)))"
# 0,1,2,3,4,20,21,22,23,24,40,...,144
```

Per-object rates from 5 episodes are noisy; use it to decide whether a full run is worth it,
not to compare checkpoints.