eval_stack: pi0.5 negmesh8 160-episode replay eval harness, job scripts, protocol docs
51c2c72 verified |
Download eval_stack/docs/EVAL_SETS.md from Ronaldo-GOAT/transfer: direct link, hf CLI and curl.
- Browser
- Download file 16.4 kB
-
https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/docs/EVAL_SETS.md
- Command line
-
hf download hf://Ronaldo-GOAT/transfer/eval_stack/docs/EVAL_SETS.md
-
curl -L -o EVAL_SETS.md https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/docs/EVAL_SETS.md
16.4 kB
| # RoboCasa eval sets for the pi0.5 ctc502 lineage — which is which, and what not to mix | |
| Written 2026-09-22. There are **three** different PickPlaceCounterToCabinet eval sets in | |
| play, they use **different objects**, **different seeds** and **different policy seeds**, | |
| and their base-checkpoint numbers are **not comparable**. Mixing them silently produces | |
| a wrong answer, so read this before quoting any number. | |
| ## The three sets | |
| | # | Set | Episodes | Env seeds | Policy / action seed | Objects | Base (60k) result | Status | | |
| |---|---|---|---|---|---|---|---| | |
| | 1 | **200-episode main** | 200 | 0–199 | 42 | the task's general object distribution | **135/200 = 67.5% — RETRACTED** | re-run required | | |
| | 2 | **160 negmesh8** | 8 obj × 20 | 10000–10019 | **42** | **our 8** (below) | **7/160 = 4.375%** | **THE SET OUR RUNS TARGET** | | |
| | 3 | 160 mimicgen8 (bundled) | 8 obj × 20 | — | **12345** | a *different* 8 (below) | 15/160 = 9.4% | actaug only — **DO NOT USE** | | |
| ### Set 2 objects — ours | |
| ``` | |
| aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023 blender_jug__BlenderJug024 | |
| jar__Jar025 juice__Juice008 syrup_bottle__SyrupBottle006 | |
| teapot__teapot_7 wine__wine_5 | |
| ``` | |
| ### Set 3 objects — NOT ours | |
| ``` | |
| donut_5 Jar023 MeasuringCup009 SoapDispenser010 | |
| steak_8 SyrupBottle005 teapot_6 teapot_7 | |
| ``` | |
| ## THE TRAP | |
| `teapot_7` is the **only** object the two 160-episode sets share. Everything else differs. | |
| Therefore: | |
| > **`actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_...` CANNOT score | |
| > our checkpoints.** It replays scenes built around seven objects our policy was never | |
| > trained on and does not contain six of ours. | |
| > **`actaug/eval/table2_exact160/eval_exact160.sh` must not be used either.** It is a | |
| > **GR00T** harness: it checks for `model.safetensors.index.json`, serves the policy through | |
| > myGR00T's `inference_service.py`, and passes `--action_horizon 16`. Our checkpoints are | |
| > **orbax/JAX pi0.5** — no safetensors, `action_horizon=50` — so it cannot even load them. | |
| > **The policy seed differs: 42 (ours, set 2) vs 12345 (set 3).** A run that quotes set 3's | |
| > seed against set 2's bank is reproducible and wrong. | |
| ## What set 1's retraction means | |
| 135/200 = 67.5% was produced by the *unseeded / history-dependent* code path in | |
| `examples/robocasa/main.py` (env built once, `env.reset()` called with no seed, so episode | |
| *i* depended on how many resets preceded it). It is **not reproducible from the seeds it | |
| listed** and has been retracted upstream. An 8-episode recheck under the fixed code gave | |
| 1/8. Any number for set 1 must be re-measured; do not cite 67.5%. | |
| Note set 2's 7/160 came from the **older** code path too. Expect our re-run to land *near* | |
| it, not on it. A wildly different result (0/160, or 60/160) means the replay branch is | |
| broken — stop and investigate rather than reporting the number. | |
| ## Two different noise conventions — check which one you want | |
| | Scheme | Formula | Where | | |
| |---|---|---| | |
| | `per_query_epid` | `default_rng(action_seed + ep_id*1_000_003 + query_idx)` per query | this repo's `examples/robocasa/main.py` (default) | | |
| | `per_episode_reset` | `default_rng(action_seed)` once per episode, one draw per query | upstream `examples/robocasa/eval_single_task_seeded.py` — **what produced set 2's 7/160** | | |
| `--args.noise_scheme per_episode_reset` selects the second. EVAL_SEEDS.md states this | |
| scheme "did not exist anywhere in the code"; that is wrong — it exists in | |
| `eval_single_task_seeded.py`, which is the harness the base numbers came from. EVAL_SEEDS.md | |
| only inspected `main.py`. | |
| ## Norm stats and action horizon — the other silent-corruption trap | |
| EVAL_SEEDS.md says to serve the ctc502 60k base checkpoint under | |
| `assets/robocasa365_human300/norm_stats.json` at `action_horizon=10`. **Both are wrong.** | |
| Upstream's real training config (`pi05_robocasa_pickplace_counter_to_cabinet_502_60k`, in | |
| `Ronaldo-GOAT/pi05-neg-mesh:training/openpi/src/openpi/training/config.py`) uses: | |
| * `action_horizon=50`, `action_dim=32`; | |
| * no `AssetsConfig` override and `repo_id=None` → `asset_id=None` → norm stats computed from | |
| the 502-demo dataset itself and written **flat** to `<ckpt>/assets/norm_stats.json`; | |
| * `use_quantile_norm` never assigned → stays `False` → **z-score**, not quantile. | |
| Measured: the checkpoint's own flat stats and this repo's `ctc502_qnorm` agree on every real | |
| dimension (state 0–15, actions 0–11) to 4.4e-4 / 1.1e-7 — same statistics. They differ only | |
| on padding dims. `robocasa365_human300` is a **different distribution** (state.mean differs | |
| by 0.86). | |
| Use config **`pi05_robocasa_ctc502_60k_eval`** for the base checkpoint (added 2026-09-22: | |
| `action_horizon=50`, `asset_id=None`, `force_zscore_norm=True`), with | |
| `assets/robocasa_ctc502/ctc502_60k_ckpt_zscore/norm_stats.json` installed as the checkpoint's | |
| flat `assets/norm_stats.json`. | |
| Checkpoints trained **in this repo** (negmesh / mimicgen arms) are different: they really were | |
| trained under `ctc502_qnorm` quantile norm, and must keep | |
| `--policy.config pi05_robocasa_target_PickPlaceCounterToCabinet_{vace_negmesh,mimicgen_aug256}`. | |
| ## Reading the VACE vs MIMICGEN head-to-head: object coverage is NOT symmetric | |
| The two arms were trained on the same 8 objects but in wildly different proportions, so an | |
| unweighted 160-episode comparison flatters VACE and buries MIMICGEN's actual strength. | |
| MIMICGEN's 256 training episodes are distributed roughly: | |
| | object | MIMICGEN train episodes | shard | note | | |
| |---|---|---|---| | |
| | `wine__wine_5` | 102 | s2 | ~40% of MIMICGEN's data | | |
| | `teapot__teapot_7` | 53 | s2 | ~21% | | |
| | `jar__Jar025` | 5 | s1 | barely seen | | |
| | `aluminum_foil__AluminumFoil006` | **0** | s0 | **zero-shot for MIMICGEN** | | |
| (VACE is 8 x 32 = 256, i.e. uniform.) | |
| Consequences, and they are the whole reason shard 2 must not be dropped: | |
| * **`s2` (episodes 120:40, teapot_7 + wine_5) carries ~60% of MIMICGEN's training data.** | |
| Evaluating only s0+s1 would delete MIMICGEN's strongest case and weight the comparison | |
| toward objects it barely saw or never saw. That is worse than not running it. | |
| * **Report three numbers, not one:** | |
| 1. **all 160** — the figure comparable to the base checkpoint's 7/160 reference; | |
| 2. **common 140** — episodes 20:160, i.e. everything except AluminumFoil006. This is the | |
| primary head-to-head, because both arms have training data for all seven of these; | |
| 3. **AluminumFoil006, 20 episodes** — broken out separately and labelled | |
| **zero-shot for MIMICGEN**, seen-32x for VACE. Never fold this into a headline average. | |
| * Per-object counts are mandatory in every report. An aggregate alone cannot distinguish | |
| "MIMICGEN is worse" from "MIMICGEN was asked about objects it never saw". | |
| `summarize_shards.py` prints all three subsets automatically. | |
| ## The mp4 round-trip: three variants, only two of them interchangeable | |
| `mp4_roundtrip` re-encodes each observation through h264/yuv420p so the policy sees the | |
| compression artifacts it was trained on. There are three implementations in the tree and | |
| they are **not** all equivalent: | |
| | variant | encoder | frames | when | pixels | | |
| |---|---|---|---|---| | |
| | `main.py` original | imageio (ffmpeg **subprocess**) | 5, take index 2 (a P-frame) | every step | baseline | | |
| | `main.py` now (default) | imageio, unchanged | 5, take index 2 | **only on replan steps** | **identical** | | |
| | `main_optimized.py` | PyAV (in-process) | 1 (an I-frame) | only on replan steps | **DIFFERENT** | | |
| Measured on a real 256x256 observation from the bank: | |
| `imageio vs PyAV -> max|d| = 49, mean|d| = 3.48, 85% of pixels differ.` | |
| So **`main_optimized.py` is not a drop-in for `main.py`** — it changes what the policy sees. | |
| Its docstring's "the action sequence is unchanged" refers only to *moving* the call into the | |
| replan branch, not to the encoder swap. Numbers from the two harnesses are not comparable. | |
| Moving the call **is** free, and is now the default in `main.py`: the round-tripped image is | |
| read only inside `if not action_plan:`, and `img` is re-derived from `obs` at the top of every | |
| iteration, so the discarded copies could never influence anything. `--args.mp4-eager` restores | |
| the old every-step behaviour for re-checking. | |
| This mattered a lot for cost. Each imageio call forks ffmpeg and takes **~479 ms**; at two per | |
| step that is ~0.97 s/step, and job 20077 episode 0 spent 268 s on 277 steps -- exactly | |
| 2 x 479 ms x 277. A 750-step episode would cost ~12 min and a 50-episode shard ~8 h, i.e. | |
| **guaranteed to blow the 4 h `eval` cap**. Doing it only on replan steps cuts that by | |
| `replan_steps` (5x here) and brings a 50-episode shard to roughly 2.5-3 h. | |
| ## `sbatch` snapshots the job script; scripts called BY PATH are read at runtime | |
| This has bitten the project twice. The rule: | |
| | what | when it is read | does a later edit reach an already-queued job? | | |
| |---|---|---| | |
| | the `.sbatch` you submit | **copied into Slurm's spool at submit time** | **NO** — the job keeps the version it was submitted with | | |
| | anything it runs by path (`main.py`, `summarize_shards.py`, `assert_norm_config.py`) | read when the job actually starts | **YES** | | |
| So a queued job runs **old job-script logic against new Python**. Consequences seen here: | |
| * Jobs 20086-20090 were submitted before the normalization assert was added to | |
| `robocasa_eval.sbatch`, so they run **without** it, while 20096+ have it. Harmless in this | |
| case (their config was already correct), but it means "I added a guard" does not imply | |
| "every job in the queue is guarded". | |
| * Nearly worse: `EXTRA_ARGS` (the passthrough that carries `--args.mp4-eager`) and the det_C | |
| submission happened in the same command. Had the order been reversed, det_C would have | |
| silently ignored the flag, run the *fast* path, and become a second fast-path run -- | |
| producing a determinism "control" that looked fine and controlled for nothing. | |
| **Verify, do not assume.** Dump what a queued job will actually execute: | |
| ```bash | |
| scontrol write batch_script <jobid> /tmp/spooled.sh | |
| grep -c EXTRA_ARGS /tmp/spooled.sh # 0 means the flag will be ignored | |
| grep -c assert_norm_config /tmp/spooled.sh | |
| ``` | |
| Corollary: changing `main.py` mid-flight DOES change every queued job. That is why the mp4 | |
| round-trip move landed cleanly (only the already-running det_A kept the old path) -- but it is | |
| also why a careless edit to `main.py` would silently split a sweep into two incomparable halves. | |
| Freeze `main.py` for the duration of a sweep; put run-to-run variation in job env vars instead. | |
| ## The `l40s` pool is heterogeneous -- asking for too many CPUs strands a job | |
| `sinfo -p l40s -N -o "%N %c %m %G"` (2026-09-22): | |
| | node(s) | CPUs | mem | GPUs | | |
| |---|---|---|---| | |
| | `l40s-dy-g6e-1x-[1-4]` | **4** | 31 GB | 1 | | |
| | `l40s-dy-g6e-2x-[1-2]` | **8** | 62 GB | 1 | | |
| | `l40s-dy-g6e-4x-[1-2]` | **16** | 121 GB | 1 | | |
| | `l40s-st-g6e-12x-1` | 48 | 365 GB | 4 | | |
| | `a100-st-p4d-cb-[1-2]` | 96 | 1.1 TB | 8 | | |
| The guide's "1x L40S 48GB per node, 8 on-demand nodes" hides this: the dynamic nodes range | |
| from **4 to 16 CPUs**. So a 1-GPU job's CPU/memory request silently decides how many nodes it | |
| can ever land on: | |
| | request | eligible | | |
| |---|---| | |
| | <= 4 CPU / 31 GB | every l40s node + a100 | | |
| | <= 8 CPU / 62 GB | 2x, 4x, 12x, a100 | | |
| | <= 16 CPU / 121 GB | 4x, 12x, a100 | | |
| | **> 16 CPU or > 121 GB** | **`l40s-st-12x` and a100 ONLY** | | |
| The parallel eval launcher asks for 24 CPU / 160 GB for K=6 workers, which lands in that last | |
| row -- jobs 20130/20131 sat `Reason=Resources` because of it, excluding 8 of the 10 l40s nodes. | |
| **Rule: for a K-worker eval job, request 16 CPU / 110 GB, not 24/160.** K=6 workers need | |
| ~1.5-2 cores each (MuJoCo physics + the ffmpeg subprocess the mp4 round-trip forks), so 16 is | |
| enough, and it keeps the `4x` nodes eligible. Going above 16 CPU or 121 GB buys nothing and | |
| costs most of the pool. | |
| ## GPU ARCHITECTURE CHANGES THE RESULT -- pin it, and never aggregate across it | |
| Measured 2026-09-22. This is a methodological finding, not an operational note. | |
| Three 8-episode runs of the **same checkpoint, seeds, config and protocol** | |
| (`pi05_ctc502_60k`, env seeds 0-7, action seed 42, `per_query_epid`): | |
| | run | job | node | GPU | mp4 path | result | | |
| |---|---|---|---|---|---| | |
| | det_A | 20077 | a100-st-p4d-cb-2 | A100-SXM4-40GB | eager | **5/8 = 62.5%** | | |
| | det_B | 20078 | a100-st-p4d-cb-2 | A100-SXM4-40GB (*different physical GPU*) | fast | **5/8 = 62.5%** | | |
| | det_C | 20090 | l40s-st-g6e-12x-1 | L40S | eager | **3/8 = 37.5%** | | |
| | comparison | what varies | bit-identical | outcome agreement | | |
| |---|---|---|---| | |
| | det_A vs det_B | process + code path, **same arch** | **4/8** | **8/8** | | |
| | det_A vs det_C | process, **arch** (same eager path) | 0/8 | 6/8 | | |
| | det_B vs det_C | process + path + **arch** | 0/8 | 6/8 | | |
| Per-episode outcomes: | |
| ``` | |
| det_A 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok | |
| det_B 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok <- identical to det_A, episode by episode | |
| det_C 0 -- 1 -- 2 ok 3 -- 4 -- 5 -- 6 ok 7 ok <- ep0 and ep4 flipped | |
| ``` | |
| **Within an architecture, reproducibility is excellent.** det_A and det_B ran on two *different | |
| physical A100s* with *different code paths* and still agreed on every outcome, 4/8 bit-identical | |
| including a 740-step rollout. | |
| **Across architectures, every single episode pair diverges at step 0, query 0** -- in the first | |
| rendered frame, before the policy has acted. Step-0 `max|d|` on the action, per episode: | |
| `9.731e-04, 4.554e-04, 9.731e-04, 3.131e-04, 1.460e-03, 5.123e-04, ~2e-03, 1.594e-03`. | |
| Both cross-architecture comparisons give *the same* values, because det_A and det_B agree with | |
| each other and both differ from det_C by the same amount. EGL rasterisation differs between | |
| A100 and L40S and the closed loop amplifies it. | |
| ### What is and is not established | |
| * **Established (strong):** architecture perturbs individual episodes -- step-0 divergence on | |
| 8 of 8 pairs. Therefore **every comparison must hold GPU architecture fixed.** | |
| * **NOT established:** that L40S scores lower than A100. n=8, and 62.5% vs 37.5% is two | |
| episodes (binomial p ~ 0.3). **Do not quote 37.5% as "the L40S number".** The honest claim is | |
| about the mechanism, not the direction. | |
| * Corollary: our numbers are **not comparable to upstream's on hardware grounds alone**, | |
| independently of the config and protocol differences already documented above. That is a | |
| further reason to re-measure the base ourselves rather than trust any published figure. | |
| ### How pinning is enforced (`-p` does NOT work) | |
| The site submit filter **rewrites every `--qos=eval` job's partition to `l40s,a100`** -- observed | |
| on jobs 20076, 20077 and 20149, where `#SBATCH -p l40s` came back as `Partition=l40s,a100`. The | |
| only way to keep a job off A100 is to exclude the nodes by name: | |
| ```bash | |
| #SBATCH --exclude=l40s-dy-g6e-1x-[1-4],l40s-dy-g6e-2x-[1-2],a100-st-p4d-cb-[1-2] | |
| ``` | |
| (the `1x`/`2x` entries are the AWS-unobtainable nodes; the `a100` entries are the pin). | |
| Verify with `scontrol show job <id> | grep ExcNodeList`, and confirm placement afterwards with | |
| `sacct -j <id> -X --format=NodeList`. | |
| ### Guard rails now in place | |
| * `examples/robocasa/main.py` records `gpu_model` (from `nvidia-smi --query-gpu=name`) **into | |
| every `stats.json`**, so the artifact is self-describing rather than relying on a job log. | |
| * `summarize_shards.py` **refuses to aggregate** shards whose `gpu_model` differs (exit 3, loud | |
| banner, per-architecture subtotals instead of a combined number). | |
| * det_A/det_B/det_C `stats.json` were backfilled with `gpu_model` plus a `gpu_model_source` | |
| field recording that it was backfilled from the job log, not measured in-harness. | |
| ## Paths | |
| | What | Where | | |
| |---|---| | |
| | negmesh8 replay bank (set 2), 160 eps, verified 968/968 files | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/episodes/` | | |
| | its ground truth | `/fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/GROUND_TRUTH.json` | | |
| | base 60k params (staged) | `/ckpt/jonghoon/eval_ckpts/pi05_ctc502_60k/` (l40s AZ) | | |
| | base 60k tar (portable) | `/s3ckpt/jonghoon/datasets/transfer_pi05_ctc502_60k.tar` (**params only — no assets/**) | | |
| | eval job | `/fsx/home/jonghoon/jobs/robocasa_eval.sbatch` | | |
| | shard summariser | `/fsx/home/jonghoon/jobs/robocasa_eval/summarize_shards.py` | | |