transfer / eval_stack /docs /EVAL_SETS.md
Ronaldo-GOAT's picture
eval_stack: pi0.5 negmesh8 160-episode replay eval harness, job scripts, protocol docs
51c2c72 verified
|
Raw History Blame Contribute Delete
16.4 kB
# RoboCasa eval sets for the pi0.5 ctc502 lineage — which is which, and what not to mix
Written 2026-09-22. There are **three** different PickPlaceCounterToCabinet eval sets in
play, they use **different objects**, **different seeds** and **different policy seeds**,
and their base-checkpoint numbers are **not comparable**. Mixing them silently produces
a wrong answer, so read this before quoting any number.
## The three sets
| # | Set | Episodes | Env seeds | Policy / action seed | Objects | Base (60k) result | Status |
|---|---|---|---|---|---|---|---|
| 1 | **200-episode main** | 200 | 0–199 | 42 | the task's general object distribution | **135/200 = 67.5% — RETRACTED** | re-run required |
| 2 | **160 negmesh8** | 8 obj × 20 | 10000–10019 | **42** | **our 8** (below) | **7/160 = 4.375%** | **THE SET OUR RUNS TARGET** |
| 3 | 160 mimicgen8 (bundled) | 8 obj × 20 | — | **12345** | a *different* 8 (below) | 15/160 = 9.4% | actaug only — **DO NOT USE** |
### Set 2 objects — ours
```
aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023 blender_jug__BlenderJug024
jar__Jar025 juice__Juice008 syrup_bottle__SyrupBottle006
teapot__teapot_7 wine__wine_5
```
### Set 3 objects — NOT ours
```
donut_5 Jar023 MeasuringCup009 SoapDispenser010
steak_8 SyrupBottle005 teapot_6 teapot_7
```
## THE TRAP
`teapot_7` is the **only** object the two 160-episode sets share. Everything else differs.
Therefore:
> **`actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_...` CANNOT score
> our checkpoints.** It replays scenes built around seven objects our policy was never
> trained on and does not contain six of ours.
> **`actaug/eval/table2_exact160/eval_exact160.sh` must not be used either.** It is a
> **GR00T** harness: it checks for `model.safetensors.index.json`, serves the policy through
> myGR00T's `inference_service.py`, and passes `--action_horizon 16`. Our checkpoints are
> **orbax/JAX pi0.5** — no safetensors, `action_horizon=50` — so it cannot even load them.
> **The policy seed differs: 42 (ours, set 2) vs 12345 (set 3).** A run that quotes set 3's
> seed against set 2's bank is reproducible and wrong.
## What set 1's retraction means
135/200 = 67.5% was produced by the *unseeded / history-dependent* code path in
`examples/robocasa/main.py` (env built once, `env.reset()` called with no seed, so episode
*i* depended on how many resets preceded it). It is **not reproducible from the seeds it
listed** and has been retracted upstream. An 8-episode recheck under the fixed code gave
1/8. Any number for set 1 must be re-measured; do not cite 67.5%.
Note set 2's 7/160 came from the **older** code path too. Expect our re-run to land *near*
it, not on it. A wildly different result (0/160, or 60/160) means the replay branch is
broken — stop and investigate rather than reporting the number.
## Two different noise conventions — check which one you want
| Scheme | Formula | Where |
|---|---|---|
| `per_query_epid` | `default_rng(action_seed + ep_id*1_000_003 + query_idx)` per query | this repo's `examples/robocasa/main.py` (default) |
| `per_episode_reset` | `default_rng(action_seed)` once per episode, one draw per query | upstream `examples/robocasa/eval_single_task_seeded.py` — **what produced set 2's 7/160** |
`--args.noise_scheme per_episode_reset` selects the second. EVAL_SEEDS.md states this
scheme "did not exist anywhere in the code"; that is wrong — it exists in
`eval_single_task_seeded.py`, which is the harness the base numbers came from. EVAL_SEEDS.md
only inspected `main.py`.
## Norm stats and action horizon — the other silent-corruption trap
EVAL_SEEDS.md says to serve the ctc502 60k base checkpoint under
`assets/robocasa365_human300/norm_stats.json` at `action_horizon=10`. **Both are wrong.**
Upstream's real training config (`pi05_robocasa_pickplace_counter_to_cabinet_502_60k`, in
`Ronaldo-GOAT/pi05-neg-mesh:training/openpi/src/openpi/training/config.py`) uses:
* `action_horizon=50`, `action_dim=32`;
* no `AssetsConfig` override and `repo_id=None` → `asset_id=None` → norm stats computed from
the 502-demo dataset itself and written **flat** to `<ckpt>/assets/norm_stats.json`;
* `use_quantile_norm` never assigned → stays `False` → **z-score**, not quantile.
Measured: the checkpoint's own flat stats and this repo's `ctc502_qnorm` agree on every real
dimension (state 0–15, actions 0–11) to 4.4e-4 / 1.1e-7 — same statistics. They differ only
on padding dims. `robocasa365_human300` is a **different distribution** (state.mean differs
by 0.86).
Use config **`pi05_robocasa_ctc502_60k_eval`** for the base checkpoint (added 2026-09-22:
`action_horizon=50`, `asset_id=None`, `force_zscore_norm=True`), with
`assets/robocasa_ctc502/ctc502_60k_ckpt_zscore/norm_stats.json` installed as the checkpoint's
flat `assets/norm_stats.json`.
Checkpoints trained **in this repo** (negmesh / mimicgen arms) are different: they really were
trained under `ctc502_qnorm` quantile norm, and must keep
`--policy.config pi05_robocasa_target_PickPlaceCounterToCabinet_{vace_negmesh,mimicgen_aug256}`.
## Reading the VACE vs MIMICGEN head-to-head: object coverage is NOT symmetric
The two arms were trained on the same 8 objects but in wildly different proportions, so an
unweighted 160-episode comparison flatters VACE and buries MIMICGEN's actual strength.
MIMICGEN's 256 training episodes are distributed roughly:
| object | MIMICGEN train episodes | shard | note |
|---|---|---|---|
| `wine__wine_5` | 102 | s2 | ~40% of MIMICGEN's data |
| `teapot__teapot_7` | 53 | s2 | ~21% |
| `jar__Jar025` | 5 | s1 | barely seen |
| `aluminum_foil__AluminumFoil006` | **0** | s0 | **zero-shot for MIMICGEN** |
(VACE is 8 x 32 = 256, i.e. uniform.)
Consequences, and they are the whole reason shard 2 must not be dropped:
* **`s2` (episodes 120:40, teapot_7 + wine_5) carries ~60% of MIMICGEN's training data.**
Evaluating only s0+s1 would delete MIMICGEN's strongest case and weight the comparison
toward objects it barely saw or never saw. That is worse than not running it.
* **Report three numbers, not one:**
1. **all 160** — the figure comparable to the base checkpoint's 7/160 reference;
2. **common 140** — episodes 20:160, i.e. everything except AluminumFoil006. This is the
primary head-to-head, because both arms have training data for all seven of these;
3. **AluminumFoil006, 20 episodes** — broken out separately and labelled
**zero-shot for MIMICGEN**, seen-32x for VACE. Never fold this into a headline average.
* Per-object counts are mandatory in every report. An aggregate alone cannot distinguish
"MIMICGEN is worse" from "MIMICGEN was asked about objects it never saw".
`summarize_shards.py` prints all three subsets automatically.
## The mp4 round-trip: three variants, only two of them interchangeable
`mp4_roundtrip` re-encodes each observation through h264/yuv420p so the policy sees the
compression artifacts it was trained on. There are three implementations in the tree and
they are **not** all equivalent:
| variant | encoder | frames | when | pixels |
|---|---|---|---|---|
| `main.py` original | imageio (ffmpeg **subprocess**) | 5, take index 2 (a P-frame) | every step | baseline |
| `main.py` now (default) | imageio, unchanged | 5, take index 2 | **only on replan steps** | **identical** |
| `main_optimized.py` | PyAV (in-process) | 1 (an I-frame) | only on replan steps | **DIFFERENT** |
Measured on a real 256x256 observation from the bank:
`imageio vs PyAV -> max|d| = 49, mean|d| = 3.48, 85% of pixels differ.`
So **`main_optimized.py` is not a drop-in for `main.py`** — it changes what the policy sees.
Its docstring's "the action sequence is unchanged" refers only to *moving* the call into the
replan branch, not to the encoder swap. Numbers from the two harnesses are not comparable.
Moving the call **is** free, and is now the default in `main.py`: the round-tripped image is
read only inside `if not action_plan:`, and `img` is re-derived from `obs` at the top of every
iteration, so the discarded copies could never influence anything. `--args.mp4-eager` restores
the old every-step behaviour for re-checking.
This mattered a lot for cost. Each imageio call forks ffmpeg and takes **~479 ms**; at two per
step that is ~0.97 s/step, and job 20077 episode 0 spent 268 s on 277 steps -- exactly
2 x 479 ms x 277. A 750-step episode would cost ~12 min and a 50-episode shard ~8 h, i.e.
**guaranteed to blow the 4 h `eval` cap**. Doing it only on replan steps cuts that by
`replan_steps` (5x here) and brings a 50-episode shard to roughly 2.5-3 h.
## `sbatch` snapshots the job script; scripts called BY PATH are read at runtime
This has bitten the project twice. The rule:
| what | when it is read | does a later edit reach an already-queued job? |
|---|---|---|
| the `.sbatch` you submit | **copied into Slurm's spool at submit time** | **NO** — the job keeps the version it was submitted with |
| anything it runs by path (`main.py`, `summarize_shards.py`, `assert_norm_config.py`) | read when the job actually starts | **YES** |
So a queued job runs **old job-script logic against new Python**. Consequences seen here:
* Jobs 20086-20090 were submitted before the normalization assert was added to
`robocasa_eval.sbatch`, so they run **without** it, while 20096+ have it. Harmless in this
case (their config was already correct), but it means "I added a guard" does not imply
"every job in the queue is guarded".
* Nearly worse: `EXTRA_ARGS` (the passthrough that carries `--args.mp4-eager`) and the det_C
submission happened in the same command. Had the order been reversed, det_C would have
silently ignored the flag, run the *fast* path, and become a second fast-path run --
producing a determinism "control" that looked fine and controlled for nothing.
**Verify, do not assume.** Dump what a queued job will actually execute:
```bash
scontrol write batch_script <jobid> /tmp/spooled.sh
grep -c EXTRA_ARGS /tmp/spooled.sh # 0 means the flag will be ignored
grep -c assert_norm_config /tmp/spooled.sh
```
Corollary: changing `main.py` mid-flight DOES change every queued job. That is why the mp4
round-trip move landed cleanly (only the already-running det_A kept the old path) -- but it is
also why a careless edit to `main.py` would silently split a sweep into two incomparable halves.
Freeze `main.py` for the duration of a sweep; put run-to-run variation in job env vars instead.
## The `l40s` pool is heterogeneous -- asking for too many CPUs strands a job
`sinfo -p l40s -N -o "%N %c %m %G"` (2026-09-22):
| node(s) | CPUs | mem | GPUs |
|---|---|---|---|
| `l40s-dy-g6e-1x-[1-4]` | **4** | 31 GB | 1 |
| `l40s-dy-g6e-2x-[1-2]` | **8** | 62 GB | 1 |
| `l40s-dy-g6e-4x-[1-2]` | **16** | 121 GB | 1 |
| `l40s-st-g6e-12x-1` | 48 | 365 GB | 4 |
| `a100-st-p4d-cb-[1-2]` | 96 | 1.1 TB | 8 |
The guide's "1x L40S 48GB per node, 8 on-demand nodes" hides this: the dynamic nodes range
from **4 to 16 CPUs**. So a 1-GPU job's CPU/memory request silently decides how many nodes it
can ever land on:
| request | eligible |
|---|---|
| <= 4 CPU / 31 GB | every l40s node + a100 |
| <= 8 CPU / 62 GB | 2x, 4x, 12x, a100 |
| <= 16 CPU / 121 GB | 4x, 12x, a100 |
| **> 16 CPU or > 121 GB** | **`l40s-st-12x` and a100 ONLY** |
The parallel eval launcher asks for 24 CPU / 160 GB for K=6 workers, which lands in that last
row -- jobs 20130/20131 sat `Reason=Resources` because of it, excluding 8 of the 10 l40s nodes.
**Rule: for a K-worker eval job, request 16 CPU / 110 GB, not 24/160.** K=6 workers need
~1.5-2 cores each (MuJoCo physics + the ffmpeg subprocess the mp4 round-trip forks), so 16 is
enough, and it keeps the `4x` nodes eligible. Going above 16 CPU or 121 GB buys nothing and
costs most of the pool.
## GPU ARCHITECTURE CHANGES THE RESULT -- pin it, and never aggregate across it
Measured 2026-09-22. This is a methodological finding, not an operational note.
Three 8-episode runs of the **same checkpoint, seeds, config and protocol**
(`pi05_ctc502_60k`, env seeds 0-7, action seed 42, `per_query_epid`):
| run | job | node | GPU | mp4 path | result |
|---|---|---|---|---|---|
| det_A | 20077 | a100-st-p4d-cb-2 | A100-SXM4-40GB | eager | **5/8 = 62.5%** |
| det_B | 20078 | a100-st-p4d-cb-2 | A100-SXM4-40GB (*different physical GPU*) | fast | **5/8 = 62.5%** |
| det_C | 20090 | l40s-st-g6e-12x-1 | L40S | eager | **3/8 = 37.5%** |
| comparison | what varies | bit-identical | outcome agreement |
|---|---|---|---|
| det_A vs det_B | process + code path, **same arch** | **4/8** | **8/8** |
| det_A vs det_C | process, **arch** (same eager path) | 0/8 | 6/8 |
| det_B vs det_C | process + path + **arch** | 0/8 | 6/8 |
Per-episode outcomes:
```
det_A 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok
det_B 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok <- identical to det_A, episode by episode
det_C 0 -- 1 -- 2 ok 3 -- 4 -- 5 -- 6 ok 7 ok <- ep0 and ep4 flipped
```
**Within an architecture, reproducibility is excellent.** det_A and det_B ran on two *different
physical A100s* with *different code paths* and still agreed on every outcome, 4/8 bit-identical
including a 740-step rollout.
**Across architectures, every single episode pair diverges at step 0, query 0** -- in the first
rendered frame, before the policy has acted. Step-0 `max|d|` on the action, per episode:
`9.731e-04, 4.554e-04, 9.731e-04, 3.131e-04, 1.460e-03, 5.123e-04, ~2e-03, 1.594e-03`.
Both cross-architecture comparisons give *the same* values, because det_A and det_B agree with
each other and both differ from det_C by the same amount. EGL rasterisation differs between
A100 and L40S and the closed loop amplifies it.
### What is and is not established
* **Established (strong):** architecture perturbs individual episodes -- step-0 divergence on
8 of 8 pairs. Therefore **every comparison must hold GPU architecture fixed.**
* **NOT established:** that L40S scores lower than A100. n=8, and 62.5% vs 37.5% is two
episodes (binomial p ~ 0.3). **Do not quote 37.5% as "the L40S number".** The honest claim is
about the mechanism, not the direction.
* Corollary: our numbers are **not comparable to upstream's on hardware grounds alone**,
independently of the config and protocol differences already documented above. That is a
further reason to re-measure the base ourselves rather than trust any published figure.
### How pinning is enforced (`-p` does NOT work)
The site submit filter **rewrites every `--qos=eval` job's partition to `l40s,a100`** -- observed
on jobs 20076, 20077 and 20149, where `#SBATCH -p l40s` came back as `Partition=l40s,a100`. The
only way to keep a job off A100 is to exclude the nodes by name:
```bash
#SBATCH --exclude=l40s-dy-g6e-1x-[1-4],l40s-dy-g6e-2x-[1-2],a100-st-p4d-cb-[1-2]
```
(the `1x`/`2x` entries are the AWS-unobtainable nodes; the `a100` entries are the pin).
Verify with `scontrol show job <id> | grep ExcNodeList`, and confirm placement afterwards with
`sacct -j <id> -X --format=NodeList`.
### Guard rails now in place
* `examples/robocasa/main.py` records `gpu_model` (from `nvidia-smi --query-gpu=name`) **into
every `stats.json`**, so the artifact is self-describing rather than relying on a job log.
* `summarize_shards.py` **refuses to aggregate** shards whose `gpu_model` differs (exit 3, loud
banner, per-architecture subtotals instead of a combined number).
* det_A/det_B/det_C `stats.json` were backfilled with `gpu_model` plus a `gpu_model_source`
field recording that it was backfilled from the job log, not measured in-harness.
## Paths
| What | Where |
|---|---|
| negmesh8 replay bank (set 2), 160 eps, verified 968/968 files | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/episodes/` |
| its ground truth | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/GROUND_TRUTH.json` |
| base 60k params (staged) | `/ckpt/jonghoon/eval_ckpts/pi05_ctc502_60k/` (l40s AZ) |
| base 60k tar (portable) | `/s3ckpt/jonghoon/datasets/transfer_pi05_ctc502_60k.tar` (**params only — no assets/**) |
| eval job | `/fsx/home/jonghoon/jobs/robocasa_eval.sbatch` |
| shard summariser | `/fsx/home/jonghoon/jobs/robocasa_eval/summarize_shards.py` |