transfer / eval_stack /README.md
Ronaldo-GOAT's picture
eval_stack: point REPLAY_ROOT at the in-repo negmesh8 bank
6f060de verified
|
Raw History Blame Contribute Delete
8.5 kB

pi0.5 RoboCasa eval stack β€” negmesh8 160-episode replay benchmark

Everything needed to reproduce our evaluation on another cluster. Built and verified 2026-09-22/23 on lab-gpu26.

What this evaluates

PickPlaceCounterToCabinet, 8 "hard mesh" objects Γ— 20 exact-replay episodes = 160. Three augmentation arms fine-tuned from the same pi05_ctc502_60k warm start, plus the base itself as reference.

Results so far (L40S, per_query_epid, action seed 42)

checkpoint all-160
base pi05_ctc502_60k 17/160 = 10.62%
published reference for the same 160 7/160 = 4.38% β€” not comparable, see below

Per object (base): AluminumFoil006 0/20 Β· BlenderJug023 3/20 Β· BlenderJug024 5/20 Β· Jar025 4/20 Β· Juice008 0/20 Β· SyrupBottle006 2/20 Β· teapot_7 2/20 Β· wine_5 1/20

THE THREE SERVING BUGS β€” read before running anything

The published 4.38% came from a harness that served the checkpoint wrongly. All three must be right or your numbers are meaningless:

wrong (published) correct
action_horizon 10 50
normalization quantile (q01/q99) z-score for the base
norm stats robocasa365_human300 the checkpoint's own ctc502 stats

Measured impact: 1/8 β†’ 5/8 on identical seeded episodes; 7/160 β†’ 17/160 on the replay set. jobs/robocasa_eval/assert_norm_config.py gates every job against this β€” run it, don't skip it.

Fine-tuned arms differ from the base here: VACE / MIMICGEN / POSE6DAUG were trained with use_quantile_norm=True on ctc502_qnorm, so they must be served quantile, not z-score. Each ships its own assets/ctc502_qnorm/. Crossing these silently corrupts results.

TWO different "VACE" checkpoint sets exist β€” do not cross them

HF repo run init data ships assets/ serve with
pi05-VACE-negmesh8-243ep-ctc502-60k-30k-h100x2 19500 pi05_ctc502_60k 243 eps (drop13) ctc502_qnorm ..._vace_negmesh (quantile)
pi05-transfer-vace-negmesh256-allintra-bs64-30k-h100x4 18808 pi05_base 256 eps robocasa365_human300 a config with asset_id="robocasa365_human300"

Only the first is the VACE arm of the three-way comparison in this package. The second is an earlier run on a different initialization, different data and different normalization; its numbers are not comparable to anything here.

Each checkpoint ships its own assets/<asset_id>/, so policy_config.py resolves from the checkpoint itself β€” serving one with the other's config fails loudly (norm stats not found ... checkpoint assets/ contains: [...]) rather than corrupting silently. Run jobs/robocasa_eval/assert_norm_config.py and trust it.

GPU ARCHITECTURE IS A CONTROLLED VARIABLE

A100 vs L40S flips ~2 of 8 episodes. Every cross-architecture episode pair diverges at step 0, query 0 β€” the initial EGL rasterisation differs. Within one architecture, two different physical GPUs agreed on 8/8 outcomes.

Run every checkpoint you intend to compare on the same GPU model. Our numbers are L40S. summarize_shards.py refuses to aggregate across mixed gpu_model values.

Setup

  1. openpi from Ronaldo-GOAT/transfer (transfer.tar) β€” includes the vendored robosuite with load_model_on_init. The PyPI robosuite 1.5.2 reports the same version string but lacks it and will fail gym.make. Uninstall PyPI robosuite, then pip install -e third_party/robosuite, and verify robosuite.__file__ is inside the bundle.

  2. RoboCasa assets (~16 GB extracted, 123,586 files) β€” jobs/robocasa_eval/fetch_robocasa_assets.py. Not bundled. robocasa's own download_kitchen_assets.py calls input() and omits the fixtures archive; use ours.

  3. The 160-episode replay bank β€” get the right one. It now ships in this repo at eval_stack/replay_bank_negmesh8/episodes/<object>/episode_0NN/ (968 files, 40.6 MB). Mirror of Ronaldo-GOAT/pi05-neg-mesh:episodes/. 8 objects x 20 episodes, env seeds 10000-10019, policy/action seed 42. Each episode ships initial_model.xml.gz, initial_sim_state.npz, initial_observation.npz, episode_metadata.pkl, replay_state.json, result.json. ~40 MB total.

    Verify you have the right bank β€” ls episodes/ must be exactly these 8:

    aluminum_foil__AluminumFoil006   blender_jug__BlenderJug023
    blender_jug__BlenderJug024       jar__Jar025
    juice__Juice008                  syrup_bottle__SyrupBottle006
    teapot__teapot_7                 wine__wine_5
    

    ep_id -> object is contiguous blocks of 20 in that alphabetical order: 0-19 AluminumFoil006, 20-39 BlenderJug023, 40-59 BlenderJug024, 60-79 Jar025, 80-99 Juice008, 100-119 SyrupBottle006, 120-139 teapot_7, 140-159 wine_5.

    There is a second, incompatible 160-episode bank shipped in the same upstream repo: actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_*. Its objects are donut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6, teapot_7 β€” only teapot_7 overlaps ours, and its policy seed is 12345, not 42. It cannot score these checkpoints, and its eval_exact160.sh is a GR00T harness that cannot even load them. Watch for the near-collisions: Jar023/Jar025, SyrupBottle005/006, teapot_6/7. See docs/EVAL_SETS.md.

  4. Overlay harness/ onto the openpi tree: main.py, get_eval_stats.py β†’ examples/robocasa/ Β· serve_policy.py β†’ scripts/ Β· policy.py, policy_config.py β†’ src/openpi/policies/ Β· config.py β†’ src/openpi/training/

Document precedence

README.md (this file) > docs/EVAL_SETS.md > docs/EVAL_SEEDS.md. The last is the upstream doc kept for provenance; it states action_horizon=10, which is wrong for every checkpoint here. Its header now says so.

Run β€” the 160-episode replay benchmark (what all our numbers use)

CKPT_DIR=<step_dir>            # containing params/ and assets/
CONFIG_NAME=<see below>
REPLAY_ROOT=eval_stack/replay_bank_negmesh8/episodes
ENV_SEED_OFFSET=0 NUM_TRIALS=160
bash jobs/robocasa_eval.sbatch
python jobs/robocasa_eval/summarize_shards.py --dir <out>

Configs: base β†’ pi05_robocasa_ctc502_60k_eval (z-score) Β· VACE β†’ ..._vace_negmesh Β· MIMICGEN β†’ ..._mimicgen_aug256 Β· POSE6DAUG β†’ ..._pose6daug_neg256 (all quantile).

Fixed protocol: action_seed=42, replan_steps=5, resize_size=224, camera 256, mp4_roundtrip=True, generative_textures=False, noise default_rng(42 + ep_id*1_000_003 + query_idx), and XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0.

Two settings that cost us hours

  • XLA_PYTHON_CLIENT_MEM_FRACTION=0.35, not 0.9. At 0.9 JAX takes 43 GB of a 48 GB L40S and the MuJoCo/EGL contexts die with GL_FRAMEBUFFER_UNSUPPORTED (0x8cdd) once you run more than ~2 workers.
  • mp4 round-trip belongs inside the if not action_plan: branch. Outside it, the encode runs every step and is discarded on 4 of 5 β€” 3.8Γ— slower, measured on identical episodes with identical outcomes.

Parallelism: one server per GPU, K=4 workers, sharded by --args.env_seed_offset. Verified bit-identical noise and 100% outcome agreement vs serial β€” verify_parallel_seeding.py. ~2.5Γ— aggregate speedup, ~2.2 h per 160-episode target.

Reproducibility

Everything the code controls is exact. The residual is MuJoCo EGL rasterisation, which is not bit-reproducible across processes (~1 pixel / 1 LSB). Measured: 4/8 episodes bit-identical across cold processes, 8/8 agreeing on outcome, identical aggregate. Long rollouts diverge more often; one 740-step episode was bit-identical throughout.

Report success rates over the full set, never per-episode outcomes.

Comparability caveats

  • MIMICGEN has 0 AluminumFoil006 training episodes (VACE has 29) and only 5 Jar025 (VACE 30). Those 40 episodes are effectively zero-shot for it. Use the common-140 subset as the head-to-head; report the 20 AluminumFoil006 episodes separately.
  • POSE6DAUG stopped at 5000 steps; VACE/MIMICGEN ran to 30000. Compare only at matched steps (1000/1500/5000), where all three saw the same LR schedule (decay_steps=30_000 deliberately unchanged).
  • 20 episodes per object β‡’ one flip is 5 points. Per-object differences of one or two successes are noise.