transfer / eval_stack /docs /EVAL_SETS.md
Ronaldo-GOAT's picture
eval_stack: pi0.5 negmesh8 160-episode replay eval harness, job scripts, protocol docs
51c2c72 verified
|
Raw History Blame Contribute Delete
16.4 kB

RoboCasa eval sets for the pi0.5 ctc502 lineage β€” which is which, and what not to mix

Written 2026-09-22. There are three different PickPlaceCounterToCabinet eval sets in play, they use different objects, different seeds and different policy seeds, and their base-checkpoint numbers are not comparable. Mixing them silently produces a wrong answer, so read this before quoting any number.

The three sets

# Set Episodes Env seeds Policy / action seed Objects Base (60k) result Status
1 200-episode main 200 0–199 42 the task's general object distribution 135/200 = 67.5% β€” RETRACTED re-run required
2 160 negmesh8 8 obj Γ— 20 10000–10019 42 our 8 (below) 7/160 = 4.375% THE SET OUR RUNS TARGET
3 160 mimicgen8 (bundled) 8 obj Γ— 20 β€” 12345 a different 8 (below) 15/160 = 9.4% actaug only β€” DO NOT USE

Set 2 objects β€” ours

aluminum_foil__AluminumFoil006   blender_jug__BlenderJug023   blender_jug__BlenderJug024
jar__Jar025                      juice__Juice008              syrup_bottle__SyrupBottle006
teapot__teapot_7                 wine__wine_5

Set 3 objects β€” NOT ours

donut_5      Jar023       MeasuringCup009   SoapDispenser010
steak_8      SyrupBottle005   teapot_6      teapot_7

THE TRAP

teapot_7 is the only object the two 160-episode sets share. Everything else differs. Therefore:

actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_... CANNOT score our checkpoints. It replays scenes built around seven objects our policy was never trained on and does not contain six of ours.

actaug/eval/table2_exact160/eval_exact160.sh must not be used either. It is a GR00T harness: it checks for model.safetensors.index.json, serves the policy through myGR00T's inference_service.py, and passes --action_horizon 16. Our checkpoints are orbax/JAX pi0.5 β€” no safetensors, action_horizon=50 β€” so it cannot even load them.

The policy seed differs: 42 (ours, set 2) vs 12345 (set 3). A run that quotes set 3's seed against set 2's bank is reproducible and wrong.

What set 1's retraction means

135/200 = 67.5% was produced by the unseeded / history-dependent code path in examples/robocasa/main.py (env built once, env.reset() called with no seed, so episode i depended on how many resets preceded it). It is not reproducible from the seeds it listed and has been retracted upstream. An 8-episode recheck under the fixed code gave 1/8. Any number for set 1 must be re-measured; do not cite 67.5%.

Note set 2's 7/160 came from the older code path too. Expect our re-run to land near it, not on it. A wildly different result (0/160, or 60/160) means the replay branch is broken β€” stop and investigate rather than reporting the number.

Two different noise conventions β€” check which one you want

Scheme Formula Where
per_query_epid default_rng(action_seed + ep_id*1_000_003 + query_idx) per query this repo's examples/robocasa/main.py (default)
per_episode_reset default_rng(action_seed) once per episode, one draw per query upstream examples/robocasa/eval_single_task_seeded.py β€” what produced set 2's 7/160

--args.noise_scheme per_episode_reset selects the second. EVAL_SEEDS.md states this scheme "did not exist anywhere in the code"; that is wrong β€” it exists in eval_single_task_seeded.py, which is the harness the base numbers came from. EVAL_SEEDS.md only inspected main.py.

Norm stats and action horizon β€” the other silent-corruption trap

EVAL_SEEDS.md says to serve the ctc502 60k base checkpoint under assets/robocasa365_human300/norm_stats.json at action_horizon=10. Both are wrong. Upstream's real training config (pi05_robocasa_pickplace_counter_to_cabinet_502_60k, in Ronaldo-GOAT/pi05-neg-mesh:training/openpi/src/openpi/training/config.py) uses:

  • action_horizon=50, action_dim=32;
  • no AssetsConfig override and repo_id=None β†’ asset_id=None β†’ norm stats computed from the 502-demo dataset itself and written flat to <ckpt>/assets/norm_stats.json;
  • use_quantile_norm never assigned β†’ stays False β†’ z-score, not quantile.

Measured: the checkpoint's own flat stats and this repo's ctc502_qnorm agree on every real dimension (state 0–15, actions 0–11) to 4.4e-4 / 1.1e-7 β€” same statistics. They differ only on padding dims. robocasa365_human300 is a different distribution (state.mean differs by 0.86).

Use config pi05_robocasa_ctc502_60k_eval for the base checkpoint (added 2026-09-22: action_horizon=50, asset_id=None, force_zscore_norm=True), with assets/robocasa_ctc502/ctc502_60k_ckpt_zscore/norm_stats.json installed as the checkpoint's flat assets/norm_stats.json.

Checkpoints trained in this repo (negmesh / mimicgen arms) are different: they really were trained under ctc502_qnorm quantile norm, and must keep --policy.config pi05_robocasa_target_PickPlaceCounterToCabinet_{vace_negmesh,mimicgen_aug256}.

Reading the VACE vs MIMICGEN head-to-head: object coverage is NOT symmetric

The two arms were trained on the same 8 objects but in wildly different proportions, so an unweighted 160-episode comparison flatters VACE and buries MIMICGEN's actual strength. MIMICGEN's 256 training episodes are distributed roughly:

object MIMICGEN train episodes shard note
wine__wine_5 102 s2 ~40% of MIMICGEN's data
teapot__teapot_7 53 s2 ~21%
jar__Jar025 5 s1 barely seen
aluminum_foil__AluminumFoil006 0 s0 zero-shot for MIMICGEN

(VACE is 8 x 32 = 256, i.e. uniform.)

Consequences, and they are the whole reason shard 2 must not be dropped:

  • s2 (episodes 120:40, teapot_7 + wine_5) carries ~60% of MIMICGEN's training data. Evaluating only s0+s1 would delete MIMICGEN's strongest case and weight the comparison toward objects it barely saw or never saw. That is worse than not running it.
  • Report three numbers, not one:
    1. all 160 β€” the figure comparable to the base checkpoint's 7/160 reference;
    2. common 140 β€” episodes 20:160, i.e. everything except AluminumFoil006. This is the primary head-to-head, because both arms have training data for all seven of these;
    3. AluminumFoil006, 20 episodes β€” broken out separately and labelled zero-shot for MIMICGEN, seen-32x for VACE. Never fold this into a headline average.
  • Per-object counts are mandatory in every report. An aggregate alone cannot distinguish "MIMICGEN is worse" from "MIMICGEN was asked about objects it never saw".

summarize_shards.py prints all three subsets automatically.

The mp4 round-trip: three variants, only two of them interchangeable

mp4_roundtrip re-encodes each observation through h264/yuv420p so the policy sees the compression artifacts it was trained on. There are three implementations in the tree and they are not all equivalent:

variant encoder frames when pixels
main.py original imageio (ffmpeg subprocess) 5, take index 2 (a P-frame) every step baseline
main.py now (default) imageio, unchanged 5, take index 2 only on replan steps identical
main_optimized.py PyAV (in-process) 1 (an I-frame) only on replan steps DIFFERENT

Measured on a real 256x256 observation from the bank: imageio vs PyAV -> max|d| = 49, mean|d| = 3.48, 85% of pixels differ. So main_optimized.py is not a drop-in for main.py β€” it changes what the policy sees. Its docstring's "the action sequence is unchanged" refers only to moving the call into the replan branch, not to the encoder swap. Numbers from the two harnesses are not comparable.

Moving the call is free, and is now the default in main.py: the round-tripped image is read only inside if not action_plan:, and img is re-derived from obs at the top of every iteration, so the discarded copies could never influence anything. --args.mp4-eager restores the old every-step behaviour for re-checking.

This mattered a lot for cost. Each imageio call forks ffmpeg and takes ~479 ms; at two per step that is ~0.97 s/step, and job 20077 episode 0 spent 268 s on 277 steps -- exactly 2 x 479 ms x 277. A 750-step episode would cost ~12 min and a 50-episode shard ~8 h, i.e. guaranteed to blow the 4 h eval cap. Doing it only on replan steps cuts that by replan_steps (5x here) and brings a 50-episode shard to roughly 2.5-3 h.

sbatch snapshots the job script; scripts called BY PATH are read at runtime

This has bitten the project twice. The rule:

what when it is read does a later edit reach an already-queued job?
the .sbatch you submit copied into Slurm's spool at submit time NO β€” the job keeps the version it was submitted with
anything it runs by path (main.py, summarize_shards.py, assert_norm_config.py) read when the job actually starts YES

So a queued job runs old job-script logic against new Python. Consequences seen here:

  • Jobs 20086-20090 were submitted before the normalization assert was added to robocasa_eval.sbatch, so they run without it, while 20096+ have it. Harmless in this case (their config was already correct), but it means "I added a guard" does not imply "every job in the queue is guarded".
  • Nearly worse: EXTRA_ARGS (the passthrough that carries --args.mp4-eager) and the det_C submission happened in the same command. Had the order been reversed, det_C would have silently ignored the flag, run the fast path, and become a second fast-path run -- producing a determinism "control" that looked fine and controlled for nothing.

Verify, do not assume. Dump what a queued job will actually execute:

scontrol write batch_script <jobid> /tmp/spooled.sh
grep -c EXTRA_ARGS /tmp/spooled.sh        # 0 means the flag will be ignored
grep -c assert_norm_config /tmp/spooled.sh

Corollary: changing main.py mid-flight DOES change every queued job. That is why the mp4 round-trip move landed cleanly (only the already-running det_A kept the old path) -- but it is also why a careless edit to main.py would silently split a sweep into two incomparable halves. Freeze main.py for the duration of a sweep; put run-to-run variation in job env vars instead.

The l40s pool is heterogeneous -- asking for too many CPUs strands a job

sinfo -p l40s -N -o "%N %c %m %G" (2026-09-22):

node(s) CPUs mem GPUs
l40s-dy-g6e-1x-[1-4] 4 31 GB 1
l40s-dy-g6e-2x-[1-2] 8 62 GB 1
l40s-dy-g6e-4x-[1-2] 16 121 GB 1
l40s-st-g6e-12x-1 48 365 GB 4
a100-st-p4d-cb-[1-2] 96 1.1 TB 8

The guide's "1x L40S 48GB per node, 8 on-demand nodes" hides this: the dynamic nodes range from 4 to 16 CPUs. So a 1-GPU job's CPU/memory request silently decides how many nodes it can ever land on:

request eligible
<= 4 CPU / 31 GB every l40s node + a100
<= 8 CPU / 62 GB 2x, 4x, 12x, a100
<= 16 CPU / 121 GB 4x, 12x, a100
> 16 CPU or > 121 GB l40s-st-12x and a100 ONLY

The parallel eval launcher asks for 24 CPU / 160 GB for K=6 workers, which lands in that last row -- jobs 20130/20131 sat Reason=Resources because of it, excluding 8 of the 10 l40s nodes.

Rule: for a K-worker eval job, request 16 CPU / 110 GB, not 24/160. K=6 workers need ~1.5-2 cores each (MuJoCo physics + the ffmpeg subprocess the mp4 round-trip forks), so 16 is enough, and it keeps the 4x nodes eligible. Going above 16 CPU or 121 GB buys nothing and costs most of the pool.

GPU ARCHITECTURE CHANGES THE RESULT -- pin it, and never aggregate across it

Measured 2026-09-22. This is a methodological finding, not an operational note.

Three 8-episode runs of the same checkpoint, seeds, config and protocol (pi05_ctc502_60k, env seeds 0-7, action seed 42, per_query_epid):

run job node GPU mp4 path result
det_A 20077 a100-st-p4d-cb-2 A100-SXM4-40GB eager 5/8 = 62.5%
det_B 20078 a100-st-p4d-cb-2 A100-SXM4-40GB (different physical GPU) fast 5/8 = 62.5%
det_C 20090 l40s-st-g6e-12x-1 L40S eager 3/8 = 37.5%
comparison what varies bit-identical outcome agreement
det_A vs det_B process + code path, same arch 4/8 8/8
det_A vs det_C process, arch (same eager path) 0/8 6/8
det_B vs det_C process + path + arch 0/8 6/8

Per-episode outcomes:

det_A  0 ok  1 --  2 ok  3 --  4 ok  5 --  6 ok  7 ok
det_B  0 ok  1 --  2 ok  3 --  4 ok  5 --  6 ok  7 ok   <- identical to det_A, episode by episode
det_C  0 --  1 --  2 ok  3 --  4 --  5 --  6 ok  7 ok   <- ep0 and ep4 flipped

Within an architecture, reproducibility is excellent. det_A and det_B ran on two different physical A100s with different code paths and still agreed on every outcome, 4/8 bit-identical including a 740-step rollout.

Across architectures, every single episode pair diverges at step 0, query 0 -- in the first rendered frame, before the policy has acted. Step-0 max|d| on the action, per episode: 9.731e-04, 4.554e-04, 9.731e-04, 3.131e-04, 1.460e-03, 5.123e-04, ~2e-03, 1.594e-03. Both cross-architecture comparisons give the same values, because det_A and det_B agree with each other and both differ from det_C by the same amount. EGL rasterisation differs between A100 and L40S and the closed loop amplifies it.

What is and is not established

  • Established (strong): architecture perturbs individual episodes -- step-0 divergence on 8 of 8 pairs. Therefore every comparison must hold GPU architecture fixed.
  • NOT established: that L40S scores lower than A100. n=8, and 62.5% vs 37.5% is two episodes (binomial p ~ 0.3). Do not quote 37.5% as "the L40S number". The honest claim is about the mechanism, not the direction.
  • Corollary: our numbers are not comparable to upstream's on hardware grounds alone, independently of the config and protocol differences already documented above. That is a further reason to re-measure the base ourselves rather than trust any published figure.

How pinning is enforced (-p does NOT work)

The site submit filter rewrites every --qos=eval job's partition to l40s,a100 -- observed on jobs 20076, 20077 and 20149, where #SBATCH -p l40s came back as Partition=l40s,a100. The only way to keep a job off A100 is to exclude the nodes by name:

#SBATCH --exclude=l40s-dy-g6e-1x-[1-4],l40s-dy-g6e-2x-[1-2],a100-st-p4d-cb-[1-2]

(the 1x/2x entries are the AWS-unobtainable nodes; the a100 entries are the pin). Verify with scontrol show job <id> | grep ExcNodeList, and confirm placement afterwards with sacct -j <id> -X --format=NodeList.

Guard rails now in place

  • examples/robocasa/main.py records gpu_model (from nvidia-smi --query-gpu=name) into every stats.json, so the artifact is self-describing rather than relying on a job log.
  • summarize_shards.py refuses to aggregate shards whose gpu_model differs (exit 3, loud banner, per-architecture subtotals instead of a combined number).
  • det_A/det_B/det_C stats.json were backfilled with gpu_model plus a gpu_model_source field recording that it was backfilled from the job log, not measured in-harness.

Paths

What Where
negmesh8 replay bank (set 2), 160 eps, verified 968/968 files /fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/episodes/
its ground truth /fsx/home/jonghoon/transfer_openpi/replay_bank_negmesh8/GROUND_TRUTH.json
base 60k params (staged) /ckpt/jonghoon/eval_ckpts/pi05_ctc502_60k/ (l40s AZ)
base 60k tar (portable) /s3ckpt/jonghoon/datasets/transfer_pi05_ctc502_60k.tar (params only β€” no assets/)
eval job /fsx/home/jonghoon/jobs/robocasa_eval.sbatch
shard summariser /fsx/home/jonghoon/jobs/robocasa_eval/summarize_shards.py