File size: 16,369 Bytes
51c2c72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
# RoboCasa eval sets for the pi0.5 ctc502 lineage β€” which is which, and what not to mix

Written 2026-09-22. There are **three** different PickPlaceCounterToCabinet eval sets in
play, they use **different objects**, **different seeds** and **different policy seeds**,
and their base-checkpoint numbers are **not comparable**. Mixing them silently produces
a wrong answer, so read this before quoting any number.

## The three sets

| # | Set | Episodes | Env seeds | Policy / action seed | Objects | Base (60k) result | Status |
|---|---|---|---|---|---|---|---|
| 1 | **200-episode main** | 200 | 0–199 | 42 | the task's general object distribution | **135/200 = 67.5% β€” RETRACTED** | re-run required |
| 2 | **160 negmesh8** | 8 obj Γ— 20 | 10000–10019 | **42** | **our 8** (below) | **7/160 = 4.375%** | **THE SET OUR RUNS TARGET** |
| 3 | 160 mimicgen8 (bundled) | 8 obj Γ— 20 | β€” | **12345** | a *different* 8 (below) | 15/160 = 9.4% | actaug only β€” **DO NOT USE** |

### Set 2 objects β€” ours
```
aluminum_foil__AluminumFoil006   blender_jug__BlenderJug023   blender_jug__BlenderJug024
jar__Jar025                      juice__Juice008              syrup_bottle__SyrupBottle006
teapot__teapot_7                 wine__wine_5
```

### Set 3 objects β€” NOT ours
```
donut_5      Jar023       MeasuringCup009   SoapDispenser010
steak_8      SyrupBottle005   teapot_6      teapot_7
```

## THE TRAP

`teapot_7` is the **only** object the two 160-episode sets share. Everything else differs.
Therefore:

> **`actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_...` CANNOT score
> our checkpoints.** It replays scenes built around seven objects our policy was never
> trained on and does not contain six of ours.

> **`actaug/eval/table2_exact160/eval_exact160.sh` must not be used either.** It is a
> **GR00T** harness: it checks for `model.safetensors.index.json`, serves the policy through
> myGR00T's `inference_service.py`, and passes `--action_horizon 16`. Our checkpoints are
> **orbax/JAX pi0.5** β€” no safetensors, `action_horizon=50` β€” so it cannot even load them.

> **The policy seed differs: 42 (ours, set 2) vs 12345 (set 3).** A run that quotes set 3's
> seed against set 2's bank is reproducible and wrong.

## What set 1's retraction means

135/200 = 67.5% was produced by the *unseeded / history-dependent* code path in
`examples/robocasa/main.py` (env built once, `env.reset()` called with no seed, so episode
*i* depended on how many resets preceded it). It is **not reproducible from the seeds it
listed** and has been retracted upstream. An 8-episode recheck under the fixed code gave
1/8. Any number for set 1 must be re-measured; do not cite 67.5%.

Note set 2's 7/160 came from the **older** code path too. Expect our re-run to land *near*
it, not on it. A wildly different result (0/160, or 60/160) means the replay branch is
broken β€” stop and investigate rather than reporting the number.

## Two different noise conventions β€” check which one you want

| Scheme | Formula | Where |
|---|---|---|
| `per_query_epid` | `default_rng(action_seed + ep_id*1_000_003 + query_idx)` per query | this repo's `examples/robocasa/main.py` (default) |
| `per_episode_reset` | `default_rng(action_seed)` once per episode, one draw per query | upstream `examples/robocasa/eval_single_task_seeded.py` β€” **what produced set 2's 7/160** |

`--args.noise_scheme per_episode_reset` selects the second. EVAL_SEEDS.md states this
scheme "did not exist anywhere in the code"; that is wrong β€” it exists in
`eval_single_task_seeded.py`, which is the harness the base numbers came from. EVAL_SEEDS.md
only inspected `main.py`.

## Norm stats and action horizon β€” the other silent-corruption trap

EVAL_SEEDS.md says to serve the ctc502 60k base checkpoint under
`assets/robocasa365_human300/norm_stats.json` at `action_horizon=10`. **Both are wrong.**
Upstream's real training config (`pi05_robocasa_pickplace_counter_to_cabinet_502_60k`, in
`Ronaldo-GOAT/pi05-neg-mesh:training/openpi/src/openpi/training/config.py`) uses:

* `action_horizon=50`, `action_dim=32`;
* no `AssetsConfig` override and `repo_id=None` β†’ `asset_id=None` β†’ norm stats computed from
  the 502-demo dataset itself and written **flat** to `<ckpt>/assets/norm_stats.json`;
* `use_quantile_norm` never assigned β†’ stays `False` β†’ **z-score**, not quantile.

Measured: the checkpoint's own flat stats and this repo's `ctc502_qnorm` agree on every real
dimension (state 0–15, actions 0–11) to 4.4e-4 / 1.1e-7 β€” same statistics. They differ only
on padding dims. `robocasa365_human300` is a **different distribution** (state.mean differs
by 0.86).

Use config **`pi05_robocasa_ctc502_60k_eval`** for the base checkpoint (added 2026-09-22:
`action_horizon=50`, `asset_id=None`, `force_zscore_norm=True`), with
`assets/robocasa_ctc502/ctc502_60k_ckpt_zscore/norm_stats.json` installed as the checkpoint's
flat `assets/norm_stats.json`.

Checkpoints trained **in this repo** (negmesh / mimicgen arms) are different: they really were
trained under `ctc502_qnorm` quantile norm, and must keep
`--policy.config pi05_robocasa_target_PickPlaceCounterToCabinet_{vace_negmesh,mimicgen_aug256}`.

## Reading the VACE vs MIMICGEN head-to-head: object coverage is NOT symmetric

The two arms were trained on the same 8 objects but in wildly different proportions, so an
unweighted 160-episode comparison flatters VACE and buries MIMICGEN's actual strength.
MIMICGEN's 256 training episodes are distributed roughly:

| object | MIMICGEN train episodes | shard | note |
|---|---|---|---|
| `wine__wine_5` | 102 | s2 | ~40% of MIMICGEN's data |
| `teapot__teapot_7` | 53 | s2 | ~21% |
| `jar__Jar025` | 5 | s1 | barely seen |
| `aluminum_foil__AluminumFoil006` | **0** | s0 | **zero-shot for MIMICGEN** |

(VACE is 8 x 32 = 256, i.e. uniform.)

Consequences, and they are the whole reason shard 2 must not be dropped:

* **`s2` (episodes 120:40, teapot_7 + wine_5) carries ~60% of MIMICGEN's training data.**
  Evaluating only s0+s1 would delete MIMICGEN's strongest case and weight the comparison
  toward objects it barely saw or never saw. That is worse than not running it.
* **Report three numbers, not one:**
  1. **all 160** β€” the figure comparable to the base checkpoint's 7/160 reference;
  2. **common 140** β€” episodes 20:160, i.e. everything except AluminumFoil006. This is the
     primary head-to-head, because both arms have training data for all seven of these;
  3. **AluminumFoil006, 20 episodes** β€” broken out separately and labelled
     **zero-shot for MIMICGEN**, seen-32x for VACE. Never fold this into a headline average.
* Per-object counts are mandatory in every report. An aggregate alone cannot distinguish
  "MIMICGEN is worse" from "MIMICGEN was asked about objects it never saw".

`summarize_shards.py` prints all three subsets automatically.

## The mp4 round-trip: three variants, only two of them interchangeable

`mp4_roundtrip` re-encodes each observation through h264/yuv420p so the policy sees the
compression artifacts it was trained on. There are three implementations in the tree and
they are **not** all equivalent:

| variant | encoder | frames | when | pixels |
|---|---|---|---|---|
| `main.py` original | imageio (ffmpeg **subprocess**) | 5, take index 2 (a P-frame) | every step | baseline |
| `main.py` now (default) | imageio, unchanged | 5, take index 2 | **only on replan steps** | **identical** |
| `main_optimized.py` | PyAV (in-process) | 1 (an I-frame) | only on replan steps | **DIFFERENT** |

Measured on a real 256x256 observation from the bank:
`imageio vs PyAV -> max|d| = 49, mean|d| = 3.48, 85% of pixels differ.`
So **`main_optimized.py` is not a drop-in for `main.py`** β€” it changes what the policy sees.
Its docstring's "the action sequence is unchanged" refers only to *moving* the call into the
replan branch, not to the encoder swap. Numbers from the two harnesses are not comparable.

Moving the call **is** free, and is now the default in `main.py`: the round-tripped image is
read only inside `if not action_plan:`, and `img` is re-derived from `obs` at the top of every
iteration, so the discarded copies could never influence anything. `--args.mp4-eager` restores
the old every-step behaviour for re-checking.

This mattered a lot for cost. Each imageio call forks ffmpeg and takes **~479 ms**; at two per
step that is ~0.97 s/step, and job 20077 episode 0 spent 268 s on 277 steps -- exactly
2 x 479 ms x 277. A 750-step episode would cost ~12 min and a 50-episode shard ~8 h, i.e.
**guaranteed to blow the 4 h `eval` cap**. Doing it only on replan steps cuts that by
`replan_steps` (5x here) and brings a 50-episode shard to roughly 2.5-3 h.

## `sbatch` snapshots the job script; scripts called BY PATH are read at runtime

This has bitten the project twice. The rule:

| what | when it is read | does a later edit reach an already-queued job? |
|---|---|---|
| the `.sbatch` you submit | **copied into Slurm's spool at submit time** | **NO** β€” the job keeps the version it was submitted with |
| anything it runs by path (`main.py`, `summarize_shards.py`, `assert_norm_config.py`) | read when the job actually starts | **YES** |

So a queued job runs **old job-script logic against new Python**. Consequences seen here:

* Jobs 20086-20090 were submitted before the normalization assert was added to
  `robocasa_eval.sbatch`, so they run **without** it, while 20096+ have it. Harmless in this
  case (their config was already correct), but it means "I added a guard" does not imply
  "every job in the queue is guarded".
* Nearly worse: `EXTRA_ARGS` (the passthrough that carries `--args.mp4-eager`) and the det_C
  submission happened in the same command. Had the order been reversed, det_C would have
  silently ignored the flag, run the *fast* path, and become a second fast-path run --
  producing a determinism "control" that looked fine and controlled for nothing.

**Verify, do not assume.** Dump what a queued job will actually execute:

```bash
scontrol write batch_script <jobid> /tmp/spooled.sh
grep -c EXTRA_ARGS /tmp/spooled.sh        # 0 means the flag will be ignored
grep -c assert_norm_config /tmp/spooled.sh
```

Corollary: changing `main.py` mid-flight DOES change every queued job. That is why the mp4
round-trip move landed cleanly (only the already-running det_A kept the old path) -- but it is
also why a careless edit to `main.py` would silently split a sweep into two incomparable halves.
Freeze `main.py` for the duration of a sweep; put run-to-run variation in job env vars instead.

## The `l40s` pool is heterogeneous -- asking for too many CPUs strands a job

`sinfo -p l40s -N -o "%N %c %m %G"` (2026-09-22):

| node(s) | CPUs | mem | GPUs |
|---|---|---|---|
| `l40s-dy-g6e-1x-[1-4]` | **4** | 31 GB | 1 |
| `l40s-dy-g6e-2x-[1-2]` | **8** | 62 GB | 1 |
| `l40s-dy-g6e-4x-[1-2]` | **16** | 121 GB | 1 |
| `l40s-st-g6e-12x-1` | 48 | 365 GB | 4 |
| `a100-st-p4d-cb-[1-2]` | 96 | 1.1 TB | 8 |

The guide's "1x L40S 48GB per node, 8 on-demand nodes" hides this: the dynamic nodes range
from **4 to 16 CPUs**. So a 1-GPU job's CPU/memory request silently decides how many nodes it
can ever land on:

| request | eligible |
|---|---|
| <= 4 CPU / 31 GB | every l40s node + a100 |
| <= 8 CPU / 62 GB | 2x, 4x, 12x, a100 |
| <= 16 CPU / 121 GB | 4x, 12x, a100 |
| **> 16 CPU or > 121 GB** | **`l40s-st-12x` and a100 ONLY** |

The parallel eval launcher asks for 24 CPU / 160 GB for K=6 workers, which lands in that last
row -- jobs 20130/20131 sat `Reason=Resources` because of it, excluding 8 of the 10 l40s nodes.

**Rule: for a K-worker eval job, request 16 CPU / 110 GB, not 24/160.** K=6 workers need
~1.5-2 cores each (MuJoCo physics + the ffmpeg subprocess the mp4 round-trip forks), so 16 is
enough, and it keeps the `4x` nodes eligible. Going above 16 CPU or 121 GB buys nothing and
costs most of the pool.

## GPU ARCHITECTURE CHANGES THE RESULT -- pin it, and never aggregate across it

Measured 2026-09-22. This is a methodological finding, not an operational note.

Three 8-episode runs of the **same checkpoint, seeds, config and protocol**
(`pi05_ctc502_60k`, env seeds 0-7, action seed 42, `per_query_epid`):

| run | job | node | GPU | mp4 path | result |
|---|---|---|---|---|---|
| det_A | 20077 | a100-st-p4d-cb-2 | A100-SXM4-40GB | eager | **5/8 = 62.5%** |
| det_B | 20078 | a100-st-p4d-cb-2 | A100-SXM4-40GB (*different physical GPU*) | fast | **5/8 = 62.5%** |
| det_C | 20090 | l40s-st-g6e-12x-1 | L40S | eager | **3/8 = 37.5%** |

| comparison | what varies | bit-identical | outcome agreement |
|---|---|---|---|
| det_A vs det_B | process + code path, **same arch** | **4/8** | **8/8** |
| det_A vs det_C | process, **arch** (same eager path) | 0/8 | 6/8 |
| det_B vs det_C | process + path + **arch** | 0/8 | 6/8 |

Per-episode outcomes:
```
det_A  0 ok  1 --  2 ok  3 --  4 ok  5 --  6 ok  7 ok
det_B  0 ok  1 --  2 ok  3 --  4 ok  5 --  6 ok  7 ok   <- identical to det_A, episode by episode
det_C  0 --  1 --  2 ok  3 --  4 --  5 --  6 ok  7 ok   <- ep0 and ep4 flipped
```

**Within an architecture, reproducibility is excellent.** det_A and det_B ran on two *different
physical A100s* with *different code paths* and still agreed on every outcome, 4/8 bit-identical
including a 740-step rollout.

**Across architectures, every single episode pair diverges at step 0, query 0** -- in the first
rendered frame, before the policy has acted. Step-0 `max|d|` on the action, per episode:
`9.731e-04, 4.554e-04, 9.731e-04, 3.131e-04, 1.460e-03, 5.123e-04, ~2e-03, 1.594e-03`.
Both cross-architecture comparisons give *the same* values, because det_A and det_B agree with
each other and both differ from det_C by the same amount. EGL rasterisation differs between
A100 and L40S and the closed loop amplifies it.

### What is and is not established

* **Established (strong):** architecture perturbs individual episodes -- step-0 divergence on
  8 of 8 pairs. Therefore **every comparison must hold GPU architecture fixed.**
* **NOT established:** that L40S scores lower than A100. n=8, and 62.5% vs 37.5% is two
  episodes (binomial p ~ 0.3). **Do not quote 37.5% as "the L40S number".** The honest claim is
  about the mechanism, not the direction.
* Corollary: our numbers are **not comparable to upstream's on hardware grounds alone**,
  independently of the config and protocol differences already documented above. That is a
  further reason to re-measure the base ourselves rather than trust any published figure.

### How pinning is enforced (`-p` does NOT work)

The site submit filter **rewrites every `--qos=eval` job's partition to `l40s,a100`** -- observed
on jobs 20076, 20077 and 20149, where `#SBATCH -p l40s` came back as `Partition=l40s,a100`. The
only way to keep a job off A100 is to exclude the nodes by name:

```bash
#SBATCH --exclude=l40s-dy-g6e-1x-[1-4],l40s-dy-g6e-2x-[1-2],a100-st-p4d-cb-[1-2]
```
(the `1x`/`2x` entries are the AWS-unobtainable nodes; the `a100` entries are the pin).
Verify with `scontrol show job <id> | grep ExcNodeList`, and confirm placement afterwards with
`sacct -j <id> -X --format=NodeList`.

### Guard rails now in place

* `examples/robocasa/main.py` records `gpu_model` (from `nvidia-smi --query-gpu=name`) **into
  every `stats.json`**, so the artifact is self-describing rather than relying on a job log.
* `summarize_shards.py` **refuses to aggregate** shards whose `gpu_model` differs (exit 3, loud
  banner, per-architecture subtotals instead of a combined number).
* det_A/det_B/det_C `stats.json` were backfilled with `gpu_model` plus a `gpu_model_source`
  field recording that it was backfilled from the job log, not measured in-harness.

## Paths

| What | Where |
|---|---|
| negmesh8 replay bank (set 2), 160 eps, verified 968/968 files | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/episodes/` |
| its ground truth | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/GROUND_TRUTH.json` |
| base 60k params (staged) | `/ckpt/jonghoon/eval_ckpts/pi05_ctc502_60k/` (l40s AZ) |
| base 60k tar (portable) | `/s3ckpt/jonghoon/datasets/transfer_pi05_ctc502_60k.tar` (**params only β€” no assets/**) |
| eval job | `/fsx/home/jonghoon/jobs/robocasa_eval.sbatch` |
| shard summariser | `/fsx/home/jonghoon/jobs/robocasa_eval/summarize_shards.py` |