# RoboCasa eval sets for the pi0.5 ctc502 lineage — which is which, and what not to mix Written 2026-09-22. There are **three** different PickPlaceCounterToCabinet eval sets in play, they use **different objects**, **different seeds** and **different policy seeds**, and their base-checkpoint numbers are **not comparable**. Mixing them silently produces a wrong answer, so read this before quoting any number. ## The three sets | # | Set | Episodes | Env seeds | Policy / action seed | Objects | Base (60k) result | Status | |---|---|---|---|---|---|---|---| | 1 | **200-episode main** | 200 | 0–199 | 42 | the task's general object distribution | **135/200 = 67.5% — RETRACTED** | re-run required | | 2 | **160 negmesh8** | 8 obj × 20 | 10000–10019 | **42** | **our 8** (below) | **7/160 = 4.375%** | **THE SET OUR RUNS TARGET** | | 3 | 160 mimicgen8 (bundled) | 8 obj × 20 | — | **12345** | a *different* 8 (below) | 15/160 = 9.4% | actaug only — **DO NOT USE** | ### Set 2 objects — ours ``` aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023 blender_jug__BlenderJug024 jar__Jar025 juice__Juice008 syrup_bottle__SyrupBottle006 teapot__teapot_7 wine__wine_5 ``` ### Set 3 objects — NOT ours ``` donut_5 Jar023 MeasuringCup009 SoapDispenser010 steak_8 SyrupBottle005 teapot_6 teapot_7 ``` ## THE TRAP `teapot_7` is the **only** object the two 160-episode sets share. Everything else differs. Therefore: > **`actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_...` CANNOT score > our checkpoints.** It replays scenes built around seven objects our policy was never > trained on and does not contain six of ours. > **`actaug/eval/table2_exact160/eval_exact160.sh` must not be used either.** It is a > **GR00T** harness: it checks for `model.safetensors.index.json`, serves the policy through > myGR00T's `inference_service.py`, and passes `--action_horizon 16`. Our checkpoints are > **orbax/JAX pi0.5** — no safetensors, `action_horizon=50` — so it cannot even load them. > **The policy seed differs: 42 (ours, set 2) vs 12345 (set 3).** A run that quotes set 3's > seed against set 2's bank is reproducible and wrong. ## What set 1's retraction means 135/200 = 67.5% was produced by the *unseeded / history-dependent* code path in `examples/robocasa/main.py` (env built once, `env.reset()` called with no seed, so episode *i* depended on how many resets preceded it). It is **not reproducible from the seeds it listed** and has been retracted upstream. An 8-episode recheck under the fixed code gave 1/8. Any number for set 1 must be re-measured; do not cite 67.5%. Note set 2's 7/160 came from the **older** code path too. Expect our re-run to land *near* it, not on it. A wildly different result (0/160, or 60/160) means the replay branch is broken — stop and investigate rather than reporting the number. ## Two different noise conventions — check which one you want | Scheme | Formula | Where | |---|---|---| | `per_query_epid` | `default_rng(action_seed + ep_id*1_000_003 + query_idx)` per query | this repo's `examples/robocasa/main.py` (default) | | `per_episode_reset` | `default_rng(action_seed)` once per episode, one draw per query | upstream `examples/robocasa/eval_single_task_seeded.py` — **what produced set 2's 7/160** | `--args.noise_scheme per_episode_reset` selects the second. EVAL_SEEDS.md states this scheme "did not exist anywhere in the code"; that is wrong — it exists in `eval_single_task_seeded.py`, which is the harness the base numbers came from. EVAL_SEEDS.md only inspected `main.py`. ## Norm stats and action horizon — the other silent-corruption trap EVAL_SEEDS.md says to serve the ctc502 60k base checkpoint under `assets/robocasa365_human300/norm_stats.json` at `action_horizon=10`. **Both are wrong.** Upstream's real training config (`pi05_robocasa_pickplace_counter_to_cabinet_502_60k`, in `Ronaldo-GOAT/pi05-neg-mesh:training/openpi/src/openpi/training/config.py`) uses: * `action_horizon=50`, `action_dim=32`; * no `AssetsConfig` override and `repo_id=None` → `asset_id=None` → norm stats computed from the 502-demo dataset itself and written **flat** to `/assets/norm_stats.json`; * `use_quantile_norm` never assigned → stays `False` → **z-score**, not quantile. Measured: the checkpoint's own flat stats and this repo's `ctc502_qnorm` agree on every real dimension (state 0–15, actions 0–11) to 4.4e-4 / 1.1e-7 — same statistics. They differ only on padding dims. `robocasa365_human300` is a **different distribution** (state.mean differs by 0.86). Use config **`pi05_robocasa_ctc502_60k_eval`** for the base checkpoint (added 2026-09-22: `action_horizon=50`, `asset_id=None`, `force_zscore_norm=True`), with `assets/robocasa_ctc502/ctc502_60k_ckpt_zscore/norm_stats.json` installed as the checkpoint's flat `assets/norm_stats.json`. Checkpoints trained **in this repo** (negmesh / mimicgen arms) are different: they really were trained under `ctc502_qnorm` quantile norm, and must keep `--policy.config pi05_robocasa_target_PickPlaceCounterToCabinet_{vace_negmesh,mimicgen_aug256}`. ## Reading the VACE vs MIMICGEN head-to-head: object coverage is NOT symmetric The two arms were trained on the same 8 objects but in wildly different proportions, so an unweighted 160-episode comparison flatters VACE and buries MIMICGEN's actual strength. MIMICGEN's 256 training episodes are distributed roughly: | object | MIMICGEN train episodes | shard | note | |---|---|---|---| | `wine__wine_5` | 102 | s2 | ~40% of MIMICGEN's data | | `teapot__teapot_7` | 53 | s2 | ~21% | | `jar__Jar025` | 5 | s1 | barely seen | | `aluminum_foil__AluminumFoil006` | **0** | s0 | **zero-shot for MIMICGEN** | (VACE is 8 x 32 = 256, i.e. uniform.) Consequences, and they are the whole reason shard 2 must not be dropped: * **`s2` (episodes 120:40, teapot_7 + wine_5) carries ~60% of MIMICGEN's training data.** Evaluating only s0+s1 would delete MIMICGEN's strongest case and weight the comparison toward objects it barely saw or never saw. That is worse than not running it. * **Report three numbers, not one:** 1. **all 160** — the figure comparable to the base checkpoint's 7/160 reference; 2. **common 140** — episodes 20:160, i.e. everything except AluminumFoil006. This is the primary head-to-head, because both arms have training data for all seven of these; 3. **AluminumFoil006, 20 episodes** — broken out separately and labelled **zero-shot for MIMICGEN**, seen-32x for VACE. Never fold this into a headline average. * Per-object counts are mandatory in every report. An aggregate alone cannot distinguish "MIMICGEN is worse" from "MIMICGEN was asked about objects it never saw". `summarize_shards.py` prints all three subsets automatically. ## The mp4 round-trip: three variants, only two of them interchangeable `mp4_roundtrip` re-encodes each observation through h264/yuv420p so the policy sees the compression artifacts it was trained on. There are three implementations in the tree and they are **not** all equivalent: | variant | encoder | frames | when | pixels | |---|---|---|---|---| | `main.py` original | imageio (ffmpeg **subprocess**) | 5, take index 2 (a P-frame) | every step | baseline | | `main.py` now (default) | imageio, unchanged | 5, take index 2 | **only on replan steps** | **identical** | | `main_optimized.py` | PyAV (in-process) | 1 (an I-frame) | only on replan steps | **DIFFERENT** | Measured on a real 256x256 observation from the bank: `imageio vs PyAV -> max|d| = 49, mean|d| = 3.48, 85% of pixels differ.` So **`main_optimized.py` is not a drop-in for `main.py`** — it changes what the policy sees. Its docstring's "the action sequence is unchanged" refers only to *moving* the call into the replan branch, not to the encoder swap. Numbers from the two harnesses are not comparable. Moving the call **is** free, and is now the default in `main.py`: the round-tripped image is read only inside `if not action_plan:`, and `img` is re-derived from `obs` at the top of every iteration, so the discarded copies could never influence anything. `--args.mp4-eager` restores the old every-step behaviour for re-checking. This mattered a lot for cost. Each imageio call forks ffmpeg and takes **~479 ms**; at two per step that is ~0.97 s/step, and job 20077 episode 0 spent 268 s on 277 steps -- exactly 2 x 479 ms x 277. A 750-step episode would cost ~12 min and a 50-episode shard ~8 h, i.e. **guaranteed to blow the 4 h `eval` cap**. Doing it only on replan steps cuts that by `replan_steps` (5x here) and brings a 50-episode shard to roughly 2.5-3 h. ## `sbatch` snapshots the job script; scripts called BY PATH are read at runtime This has bitten the project twice. The rule: | what | when it is read | does a later edit reach an already-queued job? | |---|---|---| | the `.sbatch` you submit | **copied into Slurm's spool at submit time** | **NO** — the job keeps the version it was submitted with | | anything it runs by path (`main.py`, `summarize_shards.py`, `assert_norm_config.py`) | read when the job actually starts | **YES** | So a queued job runs **old job-script logic against new Python**. Consequences seen here: * Jobs 20086-20090 were submitted before the normalization assert was added to `robocasa_eval.sbatch`, so they run **without** it, while 20096+ have it. Harmless in this case (their config was already correct), but it means "I added a guard" does not imply "every job in the queue is guarded". * Nearly worse: `EXTRA_ARGS` (the passthrough that carries `--args.mp4-eager`) and the det_C submission happened in the same command. Had the order been reversed, det_C would have silently ignored the flag, run the *fast* path, and become a second fast-path run -- producing a determinism "control" that looked fine and controlled for nothing. **Verify, do not assume.** Dump what a queued job will actually execute: ```bash scontrol write batch_script /tmp/spooled.sh grep -c EXTRA_ARGS /tmp/spooled.sh # 0 means the flag will be ignored grep -c assert_norm_config /tmp/spooled.sh ``` Corollary: changing `main.py` mid-flight DOES change every queued job. That is why the mp4 round-trip move landed cleanly (only the already-running det_A kept the old path) -- but it is also why a careless edit to `main.py` would silently split a sweep into two incomparable halves. Freeze `main.py` for the duration of a sweep; put run-to-run variation in job env vars instead. ## The `l40s` pool is heterogeneous -- asking for too many CPUs strands a job `sinfo -p l40s -N -o "%N %c %m %G"` (2026-09-22): | node(s) | CPUs | mem | GPUs | |---|---|---|---| | `l40s-dy-g6e-1x-[1-4]` | **4** | 31 GB | 1 | | `l40s-dy-g6e-2x-[1-2]` | **8** | 62 GB | 1 | | `l40s-dy-g6e-4x-[1-2]` | **16** | 121 GB | 1 | | `l40s-st-g6e-12x-1` | 48 | 365 GB | 4 | | `a100-st-p4d-cb-[1-2]` | 96 | 1.1 TB | 8 | The guide's "1x L40S 48GB per node, 8 on-demand nodes" hides this: the dynamic nodes range from **4 to 16 CPUs**. So a 1-GPU job's CPU/memory request silently decides how many nodes it can ever land on: | request | eligible | |---|---| | <= 4 CPU / 31 GB | every l40s node + a100 | | <= 8 CPU / 62 GB | 2x, 4x, 12x, a100 | | <= 16 CPU / 121 GB | 4x, 12x, a100 | | **> 16 CPU or > 121 GB** | **`l40s-st-12x` and a100 ONLY** | The parallel eval launcher asks for 24 CPU / 160 GB for K=6 workers, which lands in that last row -- jobs 20130/20131 sat `Reason=Resources` because of it, excluding 8 of the 10 l40s nodes. **Rule: for a K-worker eval job, request 16 CPU / 110 GB, not 24/160.** K=6 workers need ~1.5-2 cores each (MuJoCo physics + the ffmpeg subprocess the mp4 round-trip forks), so 16 is enough, and it keeps the `4x` nodes eligible. Going above 16 CPU or 121 GB buys nothing and costs most of the pool. ## GPU ARCHITECTURE CHANGES THE RESULT -- pin it, and never aggregate across it Measured 2026-09-22. This is a methodological finding, not an operational note. Three 8-episode runs of the **same checkpoint, seeds, config and protocol** (`pi05_ctc502_60k`, env seeds 0-7, action seed 42, `per_query_epid`): | run | job | node | GPU | mp4 path | result | |---|---|---|---|---|---| | det_A | 20077 | a100-st-p4d-cb-2 | A100-SXM4-40GB | eager | **5/8 = 62.5%** | | det_B | 20078 | a100-st-p4d-cb-2 | A100-SXM4-40GB (*different physical GPU*) | fast | **5/8 = 62.5%** | | det_C | 20090 | l40s-st-g6e-12x-1 | L40S | eager | **3/8 = 37.5%** | | comparison | what varies | bit-identical | outcome agreement | |---|---|---|---| | det_A vs det_B | process + code path, **same arch** | **4/8** | **8/8** | | det_A vs det_C | process, **arch** (same eager path) | 0/8 | 6/8 | | det_B vs det_C | process + path + **arch** | 0/8 | 6/8 | Per-episode outcomes: ``` det_A 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok det_B 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok <- identical to det_A, episode by episode det_C 0 -- 1 -- 2 ok 3 -- 4 -- 5 -- 6 ok 7 ok <- ep0 and ep4 flipped ``` **Within an architecture, reproducibility is excellent.** det_A and det_B ran on two *different physical A100s* with *different code paths* and still agreed on every outcome, 4/8 bit-identical including a 740-step rollout. **Across architectures, every single episode pair diverges at step 0, query 0** -- in the first rendered frame, before the policy has acted. Step-0 `max|d|` on the action, per episode: `9.731e-04, 4.554e-04, 9.731e-04, 3.131e-04, 1.460e-03, 5.123e-04, ~2e-03, 1.594e-03`. Both cross-architecture comparisons give *the same* values, because det_A and det_B agree with each other and both differ from det_C by the same amount. EGL rasterisation differs between A100 and L40S and the closed loop amplifies it. ### What is and is not established * **Established (strong):** architecture perturbs individual episodes -- step-0 divergence on 8 of 8 pairs. Therefore **every comparison must hold GPU architecture fixed.** * **NOT established:** that L40S scores lower than A100. n=8, and 62.5% vs 37.5% is two episodes (binomial p ~ 0.3). **Do not quote 37.5% as "the L40S number".** The honest claim is about the mechanism, not the direction. * Corollary: our numbers are **not comparable to upstream's on hardware grounds alone**, independently of the config and protocol differences already documented above. That is a further reason to re-measure the base ourselves rather than trust any published figure. ### How pinning is enforced (`-p` does NOT work) The site submit filter **rewrites every `--qos=eval` job's partition to `l40s,a100`** -- observed on jobs 20076, 20077 and 20149, where `#SBATCH -p l40s` came back as `Partition=l40s,a100`. The only way to keep a job off A100 is to exclude the nodes by name: ```bash #SBATCH --exclude=l40s-dy-g6e-1x-[1-4],l40s-dy-g6e-2x-[1-2],a100-st-p4d-cb-[1-2] ``` (the `1x`/`2x` entries are the AWS-unobtainable nodes; the `a100` entries are the pin). Verify with `scontrol show job | grep ExcNodeList`, and confirm placement afterwards with `sacct -j -X --format=NodeList`. ### Guard rails now in place * `examples/robocasa/main.py` records `gpu_model` (from `nvidia-smi --query-gpu=name`) **into every `stats.json`**, so the artifact is self-describing rather than relying on a job log. * `summarize_shards.py` **refuses to aggregate** shards whose `gpu_model` differs (exit 3, loud banner, per-architecture subtotals instead of a combined number). * det_A/det_B/det_C `stats.json` were backfilled with `gpu_model` plus a `gpu_model_source` field recording that it was backfilled from the job log, not measured in-harness. ## Paths | What | Where | |---|---| | negmesh8 replay bank (set 2), 160 eps, verified 968/968 files | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/episodes/` | | its ground truth | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/GROUND_TRUTH.json` | | base 60k params (staged) | `/ckpt/jonghoon/eval_ckpts/pi05_ctc502_60k/` (l40s AZ) | | base 60k tar (portable) | `/s3ckpt/jonghoon/datasets/transfer_pi05_ctc502_60k.tar` (**params only — no assets/**) | | eval job | `/fsx/home/jonghoon/jobs/robocasa_eval.sbatch` | | shard summariser | `/fsx/home/jonghoon/jobs/robocasa_eval/summarize_shards.py` |