Download eval_stack/docs/EVAL_SETS.md from Ronaldo-GOAT/transfer: direct link, hf CLI and curl.
- Browser
- Download file 16.4 kB
-
https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/docs/EVAL_SETS.md
- Command line
-
hf download hf://Ronaldo-GOAT/transfer/eval_stack/docs/EVAL_SETS.md
-
curl -L -o EVAL_SETS.md https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/docs/EVAL_SETS.md
RoboCasa eval sets for the pi0.5 ctc502 lineage β which is which, and what not to mix
Written 2026-09-22. There are three different PickPlaceCounterToCabinet eval sets in play, they use different objects, different seeds and different policy seeds, and their base-checkpoint numbers are not comparable. Mixing them silently produces a wrong answer, so read this before quoting any number.
The three sets
| # | Set | Episodes | Env seeds | Policy / action seed | Objects | Base (60k) result | Status |
|---|---|---|---|---|---|---|---|
| 1 | 200-episode main | 200 | 0β199 | 42 | the task's general object distribution | 135/200 = 67.5% β RETRACTED | re-run required |
| 2 | 160 negmesh8 | 8 obj Γ 20 | 10000β10019 | 42 | our 8 (below) | 7/160 = 4.375% | THE SET OUR RUNS TARGET |
| 3 | 160 mimicgen8 (bundled) | 8 obj Γ 20 | β | 12345 | a different 8 (below) | 15/160 = 9.4% | actaug only β DO NOT USE |
Set 2 objects β ours
aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023 blender_jug__BlenderJug024
jar__Jar025 juice__Juice008 syrup_bottle__SyrupBottle006
teapot__teapot_7 wine__wine_5
Set 3 objects β NOT ours
donut_5 Jar023 MeasuringCup009 SoapDispenser010
steak_8 SyrupBottle005 teapot_6 teapot_7
THE TRAP
teapot_7 is the only object the two 160-episode sets share. Everything else differs.
Therefore:
actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_...CANNOT score our checkpoints. It replays scenes built around seven objects our policy was never trained on and does not contain six of ours.
actaug/eval/table2_exact160/eval_exact160.shmust not be used either. It is a GR00T harness: it checks formodel.safetensors.index.json, serves the policy through myGR00T'sinference_service.py, and passes--action_horizon 16. Our checkpoints are orbax/JAX pi0.5 β no safetensors,action_horizon=50β so it cannot even load them.
The policy seed differs: 42 (ours, set 2) vs 12345 (set 3). A run that quotes set 3's seed against set 2's bank is reproducible and wrong.
What set 1's retraction means
135/200 = 67.5% was produced by the unseeded / history-dependent code path in
examples/robocasa/main.py (env built once, env.reset() called with no seed, so episode
i depended on how many resets preceded it). It is not reproducible from the seeds it
listed and has been retracted upstream. An 8-episode recheck under the fixed code gave
1/8. Any number for set 1 must be re-measured; do not cite 67.5%.
Note set 2's 7/160 came from the older code path too. Expect our re-run to land near it, not on it. A wildly different result (0/160, or 60/160) means the replay branch is broken β stop and investigate rather than reporting the number.
Two different noise conventions β check which one you want
| Scheme | Formula | Where |
|---|---|---|
per_query_epid |
default_rng(action_seed + ep_id*1_000_003 + query_idx) per query |
this repo's examples/robocasa/main.py (default) |
per_episode_reset |
default_rng(action_seed) once per episode, one draw per query |
upstream examples/robocasa/eval_single_task_seeded.py β what produced set 2's 7/160 |
--args.noise_scheme per_episode_reset selects the second. EVAL_SEEDS.md states this
scheme "did not exist anywhere in the code"; that is wrong β it exists in
eval_single_task_seeded.py, which is the harness the base numbers came from. EVAL_SEEDS.md
only inspected main.py.
Norm stats and action horizon β the other silent-corruption trap
EVAL_SEEDS.md says to serve the ctc502 60k base checkpoint under
assets/robocasa365_human300/norm_stats.json at action_horizon=10. Both are wrong.
Upstream's real training config (pi05_robocasa_pickplace_counter_to_cabinet_502_60k, in
Ronaldo-GOAT/pi05-neg-mesh:training/openpi/src/openpi/training/config.py) uses:
action_horizon=50,action_dim=32;- no
AssetsConfigoverride andrepo_id=Noneβasset_id=Noneβ norm stats computed from the 502-demo dataset itself and written flat to<ckpt>/assets/norm_stats.json; use_quantile_normnever assigned β staysFalseβ z-score, not quantile.
Measured: the checkpoint's own flat stats and this repo's ctc502_qnorm agree on every real
dimension (state 0β15, actions 0β11) to 4.4e-4 / 1.1e-7 β same statistics. They differ only
on padding dims. robocasa365_human300 is a different distribution (state.mean differs
by 0.86).
Use config pi05_robocasa_ctc502_60k_eval for the base checkpoint (added 2026-09-22:
action_horizon=50, asset_id=None, force_zscore_norm=True), with
assets/robocasa_ctc502/ctc502_60k_ckpt_zscore/norm_stats.json installed as the checkpoint's
flat assets/norm_stats.json.
Checkpoints trained in this repo (negmesh / mimicgen arms) are different: they really were
trained under ctc502_qnorm quantile norm, and must keep
--policy.config pi05_robocasa_target_PickPlaceCounterToCabinet_{vace_negmesh,mimicgen_aug256}.
Reading the VACE vs MIMICGEN head-to-head: object coverage is NOT symmetric
The two arms were trained on the same 8 objects but in wildly different proportions, so an unweighted 160-episode comparison flatters VACE and buries MIMICGEN's actual strength. MIMICGEN's 256 training episodes are distributed roughly:
| object | MIMICGEN train episodes | shard | note |
|---|---|---|---|
wine__wine_5 |
102 | s2 | ~40% of MIMICGEN's data |
teapot__teapot_7 |
53 | s2 | ~21% |
jar__Jar025 |
5 | s1 | barely seen |
aluminum_foil__AluminumFoil006 |
0 | s0 | zero-shot for MIMICGEN |
(VACE is 8 x 32 = 256, i.e. uniform.)
Consequences, and they are the whole reason shard 2 must not be dropped:
s2(episodes 120:40, teapot_7 + wine_5) carries ~60% of MIMICGEN's training data. Evaluating only s0+s1 would delete MIMICGEN's strongest case and weight the comparison toward objects it barely saw or never saw. That is worse than not running it.- Report three numbers, not one:
- all 160 β the figure comparable to the base checkpoint's 7/160 reference;
- common 140 β episodes 20:160, i.e. everything except AluminumFoil006. This is the primary head-to-head, because both arms have training data for all seven of these;
- AluminumFoil006, 20 episodes β broken out separately and labelled zero-shot for MIMICGEN, seen-32x for VACE. Never fold this into a headline average.
- Per-object counts are mandatory in every report. An aggregate alone cannot distinguish "MIMICGEN is worse" from "MIMICGEN was asked about objects it never saw".
summarize_shards.py prints all three subsets automatically.
The mp4 round-trip: three variants, only two of them interchangeable
mp4_roundtrip re-encodes each observation through h264/yuv420p so the policy sees the
compression artifacts it was trained on. There are three implementations in the tree and
they are not all equivalent:
| variant | encoder | frames | when | pixels |
|---|---|---|---|---|
main.py original |
imageio (ffmpeg subprocess) | 5, take index 2 (a P-frame) | every step | baseline |
main.py now (default) |
imageio, unchanged | 5, take index 2 | only on replan steps | identical |
main_optimized.py |
PyAV (in-process) | 1 (an I-frame) | only on replan steps | DIFFERENT |
Measured on a real 256x256 observation from the bank:
imageio vs PyAV -> max|d| = 49, mean|d| = 3.48, 85% of pixels differ.
So main_optimized.py is not a drop-in for main.py β it changes what the policy sees.
Its docstring's "the action sequence is unchanged" refers only to moving the call into the
replan branch, not to the encoder swap. Numbers from the two harnesses are not comparable.
Moving the call is free, and is now the default in main.py: the round-tripped image is
read only inside if not action_plan:, and img is re-derived from obs at the top of every
iteration, so the discarded copies could never influence anything. --args.mp4-eager restores
the old every-step behaviour for re-checking.
This mattered a lot for cost. Each imageio call forks ffmpeg and takes ~479 ms; at two per
step that is ~0.97 s/step, and job 20077 episode 0 spent 268 s on 277 steps -- exactly
2 x 479 ms x 277. A 750-step episode would cost ~12 min and a 50-episode shard ~8 h, i.e.
guaranteed to blow the 4 h eval cap. Doing it only on replan steps cuts that by
replan_steps (5x here) and brings a 50-episode shard to roughly 2.5-3 h.
sbatch snapshots the job script; scripts called BY PATH are read at runtime
This has bitten the project twice. The rule:
| what | when it is read | does a later edit reach an already-queued job? |
|---|---|---|
the .sbatch you submit |
copied into Slurm's spool at submit time | NO β the job keeps the version it was submitted with |
anything it runs by path (main.py, summarize_shards.py, assert_norm_config.py) |
read when the job actually starts | YES |
So a queued job runs old job-script logic against new Python. Consequences seen here:
- Jobs 20086-20090 were submitted before the normalization assert was added to
robocasa_eval.sbatch, so they run without it, while 20096+ have it. Harmless in this case (their config was already correct), but it means "I added a guard" does not imply "every job in the queue is guarded". - Nearly worse:
EXTRA_ARGS(the passthrough that carries--args.mp4-eager) and the det_C submission happened in the same command. Had the order been reversed, det_C would have silently ignored the flag, run the fast path, and become a second fast-path run -- producing a determinism "control" that looked fine and controlled for nothing.
Verify, do not assume. Dump what a queued job will actually execute:
scontrol write batch_script <jobid> /tmp/spooled.sh
grep -c EXTRA_ARGS /tmp/spooled.sh # 0 means the flag will be ignored
grep -c assert_norm_config /tmp/spooled.sh
Corollary: changing main.py mid-flight DOES change every queued job. That is why the mp4
round-trip move landed cleanly (only the already-running det_A kept the old path) -- but it is
also why a careless edit to main.py would silently split a sweep into two incomparable halves.
Freeze main.py for the duration of a sweep; put run-to-run variation in job env vars instead.
The l40s pool is heterogeneous -- asking for too many CPUs strands a job
sinfo -p l40s -N -o "%N %c %m %G" (2026-09-22):
| node(s) | CPUs | mem | GPUs |
|---|---|---|---|
l40s-dy-g6e-1x-[1-4] |
4 | 31 GB | 1 |
l40s-dy-g6e-2x-[1-2] |
8 | 62 GB | 1 |
l40s-dy-g6e-4x-[1-2] |
16 | 121 GB | 1 |
l40s-st-g6e-12x-1 |
48 | 365 GB | 4 |
a100-st-p4d-cb-[1-2] |
96 | 1.1 TB | 8 |
The guide's "1x L40S 48GB per node, 8 on-demand nodes" hides this: the dynamic nodes range from 4 to 16 CPUs. So a 1-GPU job's CPU/memory request silently decides how many nodes it can ever land on:
| request | eligible |
|---|---|
| <= 4 CPU / 31 GB | every l40s node + a100 |
| <= 8 CPU / 62 GB | 2x, 4x, 12x, a100 |
| <= 16 CPU / 121 GB | 4x, 12x, a100 |
| > 16 CPU or > 121 GB | l40s-st-12x and a100 ONLY |
The parallel eval launcher asks for 24 CPU / 160 GB for K=6 workers, which lands in that last
row -- jobs 20130/20131 sat Reason=Resources because of it, excluding 8 of the 10 l40s nodes.
Rule: for a K-worker eval job, request 16 CPU / 110 GB, not 24/160. K=6 workers need
~1.5-2 cores each (MuJoCo physics + the ffmpeg subprocess the mp4 round-trip forks), so 16 is
enough, and it keeps the 4x nodes eligible. Going above 16 CPU or 121 GB buys nothing and
costs most of the pool.
GPU ARCHITECTURE CHANGES THE RESULT -- pin it, and never aggregate across it
Measured 2026-09-22. This is a methodological finding, not an operational note.
Three 8-episode runs of the same checkpoint, seeds, config and protocol
(pi05_ctc502_60k, env seeds 0-7, action seed 42, per_query_epid):
| run | job | node | GPU | mp4 path | result |
|---|---|---|---|---|---|
| det_A | 20077 | a100-st-p4d-cb-2 | A100-SXM4-40GB | eager | 5/8 = 62.5% |
| det_B | 20078 | a100-st-p4d-cb-2 | A100-SXM4-40GB (different physical GPU) | fast | 5/8 = 62.5% |
| det_C | 20090 | l40s-st-g6e-12x-1 | L40S | eager | 3/8 = 37.5% |
| comparison | what varies | bit-identical | outcome agreement |
|---|---|---|---|
| det_A vs det_B | process + code path, same arch | 4/8 | 8/8 |
| det_A vs det_C | process, arch (same eager path) | 0/8 | 6/8 |
| det_B vs det_C | process + path + arch | 0/8 | 6/8 |
Per-episode outcomes:
det_A 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok
det_B 0 ok 1 -- 2 ok 3 -- 4 ok 5 -- 6 ok 7 ok <- identical to det_A, episode by episode
det_C 0 -- 1 -- 2 ok 3 -- 4 -- 5 -- 6 ok 7 ok <- ep0 and ep4 flipped
Within an architecture, reproducibility is excellent. det_A and det_B ran on two different physical A100s with different code paths and still agreed on every outcome, 4/8 bit-identical including a 740-step rollout.
Across architectures, every single episode pair diverges at step 0, query 0 -- in the first
rendered frame, before the policy has acted. Step-0 max|d| on the action, per episode:
9.731e-04, 4.554e-04, 9.731e-04, 3.131e-04, 1.460e-03, 5.123e-04, ~2e-03, 1.594e-03.
Both cross-architecture comparisons give the same values, because det_A and det_B agree with
each other and both differ from det_C by the same amount. EGL rasterisation differs between
A100 and L40S and the closed loop amplifies it.
What is and is not established
- Established (strong): architecture perturbs individual episodes -- step-0 divergence on 8 of 8 pairs. Therefore every comparison must hold GPU architecture fixed.
- NOT established: that L40S scores lower than A100. n=8, and 62.5% vs 37.5% is two episodes (binomial p ~ 0.3). Do not quote 37.5% as "the L40S number". The honest claim is about the mechanism, not the direction.
- Corollary: our numbers are not comparable to upstream's on hardware grounds alone, independently of the config and protocol differences already documented above. That is a further reason to re-measure the base ourselves rather than trust any published figure.
How pinning is enforced (-p does NOT work)
The site submit filter rewrites every --qos=eval job's partition to l40s,a100 -- observed
on jobs 20076, 20077 and 20149, where #SBATCH -p l40s came back as Partition=l40s,a100. The
only way to keep a job off A100 is to exclude the nodes by name:
#SBATCH --exclude=l40s-dy-g6e-1x-[1-4],l40s-dy-g6e-2x-[1-2],a100-st-p4d-cb-[1-2]
(the 1x/2x entries are the AWS-unobtainable nodes; the a100 entries are the pin).
Verify with scontrol show job <id> | grep ExcNodeList, and confirm placement afterwards with
sacct -j <id> -X --format=NodeList.
Guard rails now in place
examples/robocasa/main.pyrecordsgpu_model(fromnvidia-smi --query-gpu=name) into everystats.json, so the artifact is self-describing rather than relying on a job log.summarize_shards.pyrefuses to aggregate shards whosegpu_modeldiffers (exit 3, loud banner, per-architecture subtotals instead of a combined number).- det_A/det_B/det_C
stats.jsonwere backfilled withgpu_modelplus agpu_model_sourcefield recording that it was backfilled from the job log, not measured in-harness.
Paths
| What | Where |
|---|---|
| negmesh8 replay bank (set 2), 160 eps, verified 968/968 files | /fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/episodes/ |
| its ground truth | /fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/GROUND_TRUTH.json |
| base 60k params (staged) | /ckpt/jonghoon/eval_ckpts/pi05_ctc502_60k/ (l40s AZ) |
| base 60k tar (portable) | /s3ckpt/jonghoon/datasets/transfer_pi05_ctc502_60k.tar (params only β no assets/) |
| eval job | /fsx/home/jonghoon/jobs/robocasa_eval.sbatch |
| shard summariser | /fsx/home/jonghoon/jobs/robocasa_eval/summarize_shards.py |