File size: 16,369 Bytes
51c2c72 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 | # RoboCasa eval sets for the pi0.5 ctc502 lineage β which is which, and what not to mix
Written 2026-09-22. There are **three** different PickPlaceCounterToCabinet eval sets in
play, they use **different objects**, **different seeds** and **different policy seeds**,
and their base-checkpoint numbers are **not comparable**. Mixing them silently produces
a wrong answer, so read this before quoting any number.
## The three sets
| # | Set | Episodes | Env seeds | Policy / action seed | Objects | Base (60k) result | Status |
|---|---|---|---|---|---|---|---|
| 1 | **200-episode main** | 200 | 0β199 | 42 | the task's general object distribution | **135/200 = 67.5% β RETRACTED** | re-run required |
| 2 | **160 negmesh8** | 8 obj Γ 20 | 10000β10019 | **42** | **our 8** (below) | **7/160 = 4.375%** | **THE SET OUR RUNS TARGET** |
| 3 | 160 mimicgen8 (bundled) | 8 obj Γ 20 | β | **12345** | a *different* 8 (below) | 15/160 = 9.4% | actaug only β **DO NOT USE** |
### Set 2 objects β ours
```
aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023 blender_jug__BlenderJug024
jar__Jar025 juice__Juice008 syrup_bottle__SyrupBottle006
teapot__teapot_7 wine__wine_5
```
### Set 3 objects β NOT ours
```
donut_5 Jar023 MeasuringCup009 SoapDispenser010
steak_8 SyrupBottle005 teapot_6 teapot_7
```
## THE TRAP
`teapot_7` is the **only** object the two 160-episode sets share. Everything else differs.
Therefore:
> **`actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_...` CANNOT score
> our checkpoints.** It replays scenes built around seven objects our policy was never
> trained on and does not contain six of ours.
> **`actaug/eval/table2_exact160/eval_exact160.sh` must not be used either.** It is a
> **GR00T** harness: it checks for `model.safetensors.index.json`, serves the policy through
> myGR00T's `inference_service.py`, and passes `--action_horizon 16`. Our checkpoints are
> **orbax/JAX pi0.5** β no safetensors, `action_horizon=50` β so it cannot even load them.
> **The policy seed differs: 42 (ours, set 2) vs 12345 (set 3).** A run that quotes set 3's
> seed against set 2's bank is reproducible and wrong.
## What set 1's retraction means
135/200 = 67.5% was produced by the *unseeded / history-dependent* code path in
`examples/robocasa/main.py` (env built once, `env.reset()` called with no seed, so episode
*i* depended on how many resets preceded it). It is **not reproducible from the seeds it
listed** and has been retracted upstream. An 8-episode recheck under the fixed code gave
1/8. Any number for set 1 must be re-measured; do not cite 67.5%.
Note set 2's 7/160 came from the **older** code path too. Expect our re-run to land *near*
it, not on it. A wildly different result (0/160, or 60/160) means the replay branch is
broken β stop and investigate rather than reporting the number.
## Two different noise conventions β check which one you want
| Scheme | Formula | Where |
|---|---|---|
| `per_query_epid` | `default_rng(action_seed + ep_id*1_000_003 + query_idx)` per query | this repo's `examples/robocasa/main.py` (default) |
| `per_episode_reset` | `default_rng(action_seed)` once per episode, one draw per query | upstream `examples/robocasa/eval_single_task_seeded.py` β **what produced set 2's 7/160** |
`--args.noise_scheme per_episode_reset` selects the second. EVAL_SEEDS.md states this
scheme "did not exist anywhere in the code"; that is wrong β it exists in
`eval_single_task_seeded.py`, which is the harness the base numbers came from. EVAL_SEEDS.md
only inspected `main.py`.
## Norm stats and action horizon β the other silent-corruption trap
EVAL_SEEDS.md says to serve the ctc502 60k base checkpoint under
`assets/robocasa365_human300/norm_stats.json` at `action_horizon=10`. **Both are wrong.**
Upstream's real training config (`pi05_robocasa_pickplace_counter_to_cabinet_502_60k`, in
`Ronaldo-GOAT/pi05-neg-mesh:training/openpi/src/openpi/training/config.py`) uses:
* `action_horizon=50`, `action_dim=32`;
* no `AssetsConfig` override and `repo_id=None` β `asset_id=None` β norm stats computed from
the 502-demo dataset itself and written **flat** to `<ckpt>/assets/norm_stats.json`;
* `use_quantile_norm` never assigned β stays `False` β **z-score**, not quantile.
Measured: the checkpoint's own flat stats and this repo's `ctc502_qnorm` agree on every real
dimension (state 0β15, actions 0β11) to 4.4e-4 / 1.1e-7 β same statistics. They differ only
on padding dims. `robocasa365_human300` is a **different distribution** (state.mean differs
by 0.86).
Use config **`pi05_robocasa_ctc502_60k_eval`** for the base checkpoint (added 2026-09-22:
`action_horizon=50`, `asset_id=None`, `force_zscore_norm=True`), with
`assets/robocasa_ctc502/ctc502_60k_ckpt_zscore/norm_stats.json` installed as the checkpoint's
flat `assets/norm_stats.json`.
Checkpoints trained **in this repo** (negmesh / mimicgen arms) are different: they really were
trained under `ctc502_qnorm` quantile norm, and must keep
`--policy.config pi05_robocasa_target_PickPlaceCounterToCabinet_{vace_negmesh,mimicgen_aug256}`.
## Reading the VACE vs MIMICGEN head-to-head: object coverage is NOT symmetric
The two arms were trained on the same 8 objects but in wildly different proportions, so an
unweighted 160-episode comparison flatters VACE and buries MIMICGEN's actual strength.
MIMICGEN's 256 training episodes are distributed roughly:
| object | MIMICGEN train episodes | shard | note |
|---|---|---|---|
| `wine__wine_5` | 102 | s2 | ~40% of MIMICGEN's data |
| `teapot__teapot_7` | 53 | s2 | ~21% |
| `jar__Jar025` | 5 | s1 | barely seen |
| `aluminum_foil__AluminumFoil006` | **0** | s0 | **zero-shot for MIMICGEN** |
(VACE is 8 x 32 = 256, i.e. uniform.)
Consequences, and they are the whole reason shard 2 must not be dropped:
* **`s2` (episodes 120:40, teapot_7 + wine_5) carries ~60% of MIMICGEN's training data.**
Evaluating only s0+s1 would delete MIMICGEN's strongest case and weight the comparison
toward objects it barely saw or never saw. That is worse than not running it.
* **Report three numbers, not one:**
1. **all 160** β the figure comparable to the base checkpoint's 7/160 reference;
2. **common 140** β episodes 20:160, i.e. everything except AluminumFoil006. This is the
primary head-to-head, because both arms have training data for all seven of these;
3. **AluminumFoil006, 20 episodes** β broken out separately and labelled
**zero-shot for MIMICGEN**, seen-32x for VACE. Never fold this into a headline average.
* Per-object counts are mandatory in every report. An aggregate alone cannot distinguish
"MIMICGEN is worse" from "MIMICGEN was asked about objects it never saw".
`summarize_shards.py` prints all three subsets automatically.
## The mp4 round-trip: three variants, only two of them interchangeable
`mp4_roundtrip` re-encodes each observation through h264/yuv420p so the policy sees the
compression artifacts it was trained on. There are three implementations in the tree and
they are **not** all equivalent:
| variant | encoder | frames | when | pixels |
|---|---|---|---|---|
| `main.py` original | imageio (ffmpeg **subprocess**) | 5, take index 2 (a P-frame) | every step | baseline |
| `main.py` now (default) | imageio, unchanged | 5, take index 2 | **only on replan steps** | **identical** |
| `main_optimized.py` | PyAV (in-process) | 1 (an I-frame) | only on replan steps | **DIFFERENT** |
Measured on a real 256x256 observation from the bank:
`imageio vs PyAV -> max|d| = 49, mean|d| = 3.48, 85% of pixels differ.`
So **`main_optimized.py` is not a drop-in for `main.py`** β it changes what the policy sees.
Its docstring's "the action sequence is unchanged" refers only to *moving* the call into the
replan branch, not to the encoder swap. Numbers from the two harnesses are not comparable.
Moving the call **is** free, and is now the default in `main.py`: the round-tripped image is
read only inside `if not action_plan:`, and `img` is re-derived from `obs` at the top of every
iteration, so the discarded copies could never influence anything. `--args.mp4-eager` restores
the old every-step behaviour for re-checking.
This mattered a lot for cost. Each imageio call forks ffmpeg and takes **~479 ms**; at two per
step that is ~0.97 s/step, and job 20077 episode 0 spent 268 s on 277 steps -- exactly
2 x 479 ms x 277. A 750-step episode would cost ~12 min and a 50-episode shard ~8 h, i.e.
**guaranteed to blow the 4 h `eval` cap**. Doing it only on replan steps cuts that by
`replan_steps` (5x here) and brings a 50-episode shard to roughly 2.5-3 h.
## `sbatch` snapshots the job script; scripts called BY PATH are read at runtime
This has bitten the project twice. The rule:
| what | when it is read | does a later edit reach an already-queued job? |
|---|---|---|
| the `.sbatch` you submit | **copied into Slurm's spool at submit time** | **NO** β the job keeps the version it was submitted with |
| anything it runs by path (`main.py`, `summarize_shards.py`, `assert_norm_config.py`) | read when the job actually starts | **YES** |
So a queued job runs **old job-script logic against new Python**. Consequences seen here:
* Jobs 20086-20090 were submitted before the normalization assert was added to
`robocasa_eval.sbatch`, so they run **without** it, while 20096+ have it. Harmless in this
case (their config was already correct), but it means "I added a guard" does not imply
"every job in the queue is guarded".
* Nearly worse: `EXTRA_ARGS` (the passthrough that carries `--args.mp4-eager`) and the det_C
submission happened in the same command. Had the order been reversed, det_C would have
silently ignored the flag, run the *fast* path, and become a second fast-path run --
producing a determinism "control" that looked fine and controlled for nothing.
**Verify, do not assume.** Dump what a queued job will actually execute:
```bash
scontrol write batch_script <jobid> /tmp/spooled.sh
grep -c EXTRA_ARGS /tmp/spooled.sh # 0 means the flag will be ignored
grep -c assert_norm_config /tmp/spooled.sh
```
Corollary: changing `main.py` mid-flight DOES change every queued job. That is why the mp4
round-trip move landed cleanly (only the already-running det_A kept the old path) -- but it is
also why a careless edit to `main.py` would silently split a sweep into two incomparable halves.
Freeze `main.py` for the duration of a sweep; put run-to-run variation in job env vars instead.
## The `l40s` pool is heterogeneous -- asking for too many CPUs strands a job
`sinfo -p l40s -N -o "%N %c %m %G"` (2026-09-22):
| node(s) | CPUs | mem | GPUs |
|---|---|---|---|
| `l40s-dy-g6e-1x-[1-4]` | **4** | 31 GB | 1 |
| `l40s-dy-g6e-2x-[1-2]` | **8** | 62 GB | 1 |
| `l40s-dy-g6e-4x-[1-2]` | **16** | 121 GB | 1 |
| `l40s-st-g6e-12x-1` | 48 | 365 GB | 4 |
| `a100-st-p4d-cb-[1-2]` | 96 | 1.1 TB | 8 |
The guide's "1x L40S 48GB per node, 8 on-demand nodes" hides this: the dynamic nodes range
from **4 to 16 CPUs**. So a 1-GPU job's CPU/memory request silently decides how many nodes it
can ever land on:
| request | eligible |
|---|---|
| <= 4 CPU / 31 GB | every l40s node + a100 |
| <= 8 CPU / 62 GB | 2x, 4x, 12x, a100 |
| <= 16 CPU / 121 GB | 4x, 12x, a100 |
| **> 16 CPU or > 121 GB** | **`l40s-st-12x` and a100 ONLY** |
The parallel eval launcher asks for 24 CPU / 160 GB for K=6 workers, which lands in that last
row -- jobs 20130/20131 sat `Reason=Resources` because of it, excluding 8 of the 10 l40s nodes.
**Rule: for a K-worker eval job, request 16 CPU / 110 GB, not 24/160.** K=6 workers need
~1.5-2 cores each (MuJoCo physics + the ffmpeg subprocess the mp4 round-trip forks), so 16 is
enough, and it keeps the `4x` nodes eligible. Going above 16 CPU or 121 GB buys nothing and
costs most of the pool.
## GPU ARCHITECTURE CHANGES THE RESULT -- pin it, and never aggregate across it
Measured 2026-09-22. This is a methodological finding, not an operational note.
Three 8-episode runs of the **same checkpoint, seeds, config and protocol**
(`pi05_ctc502_60k`, env seeds 0-7, action seed 42, `per_query_epid`):
| run | job | node | GPU | mp4 path | result |
|---|---|---|---|---|---|
| det_A | 20077 | a100-st-p4d-cb-2 | A100-SXM4-40GB | eager | **5/8 = 62.5%** |
| det_B | 20078 | a100-st-p4d-cb-2 | A100-SXM4-40GB (*different physical GPU*) | fast | **5/8 = 62.5%** |
| det_C | 20090 | l40s-st-g6e-12x-1 | L40S | eager | **3/8 = 37.5%** |
| comparison | what varies | bit-identical | outcome agreement |
|---|---|---|---|
| det_A vs det_B | process + code path, **same arch** | **4/8** | **8/8** |
| det_A vs det_C | process, **arch** (same eager path) | 0/8 | 6/8 |
| det_B vs det_C | process + path + **arch** | 0/8 | 6/8 |
Per-episode outcomes:
```
det_A 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok
det_B 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok <- identical to det_A, episode by episode
det_C 0 -- 1 -- 2 ok 3 -- 4 -- 5 -- 6 ok 7 ok <- ep0 and ep4 flipped
```
**Within an architecture, reproducibility is excellent.** det_A and det_B ran on two *different
physical A100s* with *different code paths* and still agreed on every outcome, 4/8 bit-identical
including a 740-step rollout.
**Across architectures, every single episode pair diverges at step 0, query 0** -- in the first
rendered frame, before the policy has acted. Step-0 `max|d|` on the action, per episode:
`9.731e-04, 4.554e-04, 9.731e-04, 3.131e-04, 1.460e-03, 5.123e-04, ~2e-03, 1.594e-03`.
Both cross-architecture comparisons give *the same* values, because det_A and det_B agree with
each other and both differ from det_C by the same amount. EGL rasterisation differs between
A100 and L40S and the closed loop amplifies it.
### What is and is not established
* **Established (strong):** architecture perturbs individual episodes -- step-0 divergence on
8 of 8 pairs. Therefore **every comparison must hold GPU architecture fixed.**
* **NOT established:** that L40S scores lower than A100. n=8, and 62.5% vs 37.5% is two
episodes (binomial p ~ 0.3). **Do not quote 37.5% as "the L40S number".** The honest claim is
about the mechanism, not the direction.
* Corollary: our numbers are **not comparable to upstream's on hardware grounds alone**,
independently of the config and protocol differences already documented above. That is a
further reason to re-measure the base ourselves rather than trust any published figure.
### How pinning is enforced (`-p` does NOT work)
The site submit filter **rewrites every `--qos=eval` job's partition to `l40s,a100`** -- observed
on jobs 20076, 20077 and 20149, where `#SBATCH -p l40s` came back as `Partition=l40s,a100`. The
only way to keep a job off A100 is to exclude the nodes by name:
```bash
#SBATCH --exclude=l40s-dy-g6e-1x-[1-4],l40s-dy-g6e-2x-[1-2],a100-st-p4d-cb-[1-2]
```
(the `1x`/`2x` entries are the AWS-unobtainable nodes; the `a100` entries are the pin).
Verify with `scontrol show job <id> | grep ExcNodeList`, and confirm placement afterwards with
`sacct -j <id> -X --format=NodeList`.
### Guard rails now in place
* `examples/robocasa/main.py` records `gpu_model` (from `nvidia-smi --query-gpu=name`) **into
every `stats.json`**, so the artifact is self-describing rather than relying on a job log.
* `summarize_shards.py` **refuses to aggregate** shards whose `gpu_model` differs (exit 3, loud
banner, per-architecture subtotals instead of a combined number).
* det_A/det_B/det_C `stats.json` were backfilled with `gpu_model` plus a `gpu_model_source`
field recording that it was backfilled from the job log, not measured in-harness.
## Paths
| What | Where |
|---|---|
| negmesh8 replay bank (set 2), 160 eps, verified 968/968 files | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/episodes/` |
| its ground truth | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/GROUND_TRUTH.json` |
| base 60k params (staged) | `/ckpt/jonghoon/eval_ckpts/pi05_ctc502_60k/` (l40s AZ) |
| base 60k tar (portable) | `/s3ckpt/jonghoon/datasets/transfer_pi05_ctc502_60k.tar` (**params only β no assets/**) |
| eval job | `/fsx/home/jonghoon/jobs/robocasa_eval.sbatch` |
| shard summariser | `/fsx/home/jonghoon/jobs/robocasa_eval/summarize_shards.py` |
|