Download eval_stack/README.md from Ronaldo-GOAT/transfer: direct link, hf CLI and curl.
- Browser
- Download file 8.5 kB
-
https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/README.md
- Command line
-
hf download hf://Ronaldo-GOAT/transfer/eval_stack/README.md
-
curl -L -o README.md https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/README.md
pi0.5 RoboCasa eval stack β negmesh8 160-episode replay benchmark
Everything needed to reproduce our evaluation on another cluster. Built and verified 2026-09-22/23 on lab-gpu26.
What this evaluates
PickPlaceCounterToCabinet, 8 "hard mesh" objects Γ 20 exact-replay episodes = 160.
Three augmentation arms fine-tuned from the same pi05_ctc502_60k warm start, plus the
base itself as reference.
Results so far (L40S, per_query_epid, action seed 42)
| checkpoint | all-160 |
|---|---|
base pi05_ctc502_60k |
17/160 = 10.62% |
| published reference for the same 160 | 7/160 = 4.38% β not comparable, see below |
Per object (base): AluminumFoil006 0/20 Β· BlenderJug023 3/20 Β· BlenderJug024 5/20 Β· Jar025 4/20 Β· Juice008 0/20 Β· SyrupBottle006 2/20 Β· teapot_7 2/20 Β· wine_5 1/20
THE THREE SERVING BUGS β read before running anything
The published 4.38% came from a harness that served the checkpoint wrongly. All three must be right or your numbers are meaningless:
| wrong (published) | correct | |
|---|---|---|
action_horizon |
10 | 50 |
| normalization | quantile (q01/q99) | z-score for the base |
| norm stats | robocasa365_human300 |
the checkpoint's own ctc502 stats |
Measured impact: 1/8 β 5/8 on identical seeded episodes; 7/160 β 17/160 on the
replay set. jobs/robocasa_eval/assert_norm_config.py gates every job against this β
run it, don't skip it.
Fine-tuned arms differ from the base here: VACE / MIMICGEN / POSE6DAUG were trained
with use_quantile_norm=True on ctc502_qnorm, so they must be served quantile,
not z-score. Each ships its own assets/ctc502_qnorm/. Crossing these silently corrupts
results.
TWO different "VACE" checkpoint sets exist β do not cross them
| HF repo | run | init | data | ships assets/ |
serve with |
|---|---|---|---|---|---|
pi05-VACE-negmesh8-243ep-ctc502-60k-30k-h100x2 |
19500 | pi05_ctc502_60k |
243 eps (drop13) | ctc502_qnorm |
..._vace_negmesh (quantile) |
pi05-transfer-vace-negmesh256-allintra-bs64-30k-h100x4 |
18808 | pi05_base |
256 eps | robocasa365_human300 |
a config with asset_id="robocasa365_human300" |
Only the first is the VACE arm of the three-way comparison in this package. The second is an earlier run on a different initialization, different data and different normalization; its numbers are not comparable to anything here.
Each checkpoint ships its own assets/<asset_id>/, so policy_config.py resolves from the
checkpoint itself β serving one with the other's config fails loudly (norm stats not found ... checkpoint assets/ contains: [...]) rather than corrupting silently. Run
jobs/robocasa_eval/assert_norm_config.py and trust it.
GPU ARCHITECTURE IS A CONTROLLED VARIABLE
A100 vs L40S flips ~2 of 8 episodes. Every cross-architecture episode pair diverges at step 0, query 0 β the initial EGL rasterisation differs. Within one architecture, two different physical GPUs agreed on 8/8 outcomes.
Run every checkpoint you intend to compare on the same GPU model. Our numbers are
L40S. summarize_shards.py refuses to aggregate across mixed gpu_model values.
Setup
openpifromRonaldo-GOAT/transfer(transfer.tar) β includes the vendored robosuite withload_model_on_init. The PyPI robosuite 1.5.2 reports the same version string but lacks it and will failgym.make. Uninstall PyPI robosuite, thenpip install -e third_party/robosuite, and verifyrobosuite.__file__is inside the bundle.RoboCasa assets (~16 GB extracted, 123,586 files) β
jobs/robocasa_eval/fetch_robocasa_assets.py. Not bundled. robocasa's owndownload_kitchen_assets.pycallsinput()and omits the fixtures archive; use ours.The 160-episode replay bank β get the right one. It now ships in this repo at
eval_stack/replay_bank_negmesh8/episodes/<object>/episode_0NN/(968 files, 40.6 MB). Mirror ofRonaldo-GOAT/pi05-neg-mesh:episodes/. 8 objects x 20 episodes, env seeds 10000-10019, policy/action seed 42. Each episode shipsinitial_model.xml.gz,initial_sim_state.npz,initial_observation.npz,episode_metadata.pkl,replay_state.json,result.json. ~40 MB total.Verify you have the right bank β
ls episodes/must be exactly these 8:aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023 blender_jug__BlenderJug024 jar__Jar025 juice__Juice008 syrup_bottle__SyrupBottle006 teapot__teapot_7 wine__wine_5ep_id -> object is contiguous blocks of 20 in that alphabetical order: 0-19 AluminumFoil006, 20-39 BlenderJug023, 40-59 BlenderJug024, 60-79 Jar025, 80-99 Juice008, 100-119 SyrupBottle006, 120-139 teapot_7, 140-159 wine_5.
There is a second, incompatible 160-episode bank shipped in the same upstream repo:
actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_*. Its objects aredonut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6, teapot_7β onlyteapot_7overlaps ours, and its policy seed is 12345, not 42. It cannot score these checkpoints, and itseval_exact160.shis a GR00T harness that cannot even load them. Watch for the near-collisions: Jar023/Jar025, SyrupBottle005/006, teapot_6/7. Seedocs/EVAL_SETS.md.Overlay
harness/onto the openpi tree:main.py,get_eval_stats.pyβexamples/robocasa/Β·serve_policy.pyβscripts/Β·policy.py,policy_config.pyβsrc/openpi/policies/Β·config.pyβsrc/openpi/training/
Document precedence
README.md (this file) > docs/EVAL_SETS.md > docs/EVAL_SEEDS.md. The last is the
upstream doc kept for provenance; it states action_horizon=10, which is wrong for
every checkpoint here. Its header now says so.
Run β the 160-episode replay benchmark (what all our numbers use)
CKPT_DIR=<step_dir> # containing params/ and assets/
CONFIG_NAME=<see below>
REPLAY_ROOT=eval_stack/replay_bank_negmesh8/episodes
ENV_SEED_OFFSET=0 NUM_TRIALS=160
bash jobs/robocasa_eval.sbatch
python jobs/robocasa_eval/summarize_shards.py --dir <out>
Configs: base β pi05_robocasa_ctc502_60k_eval (z-score) Β· VACE β
..._vace_negmesh Β· MIMICGEN β ..._mimicgen_aug256 Β· POSE6DAUG β ..._pose6daug_neg256
(all quantile).
Fixed protocol: action_seed=42, replan_steps=5, resize_size=224, camera 256,
mp4_roundtrip=True, generative_textures=False, noise
default_rng(42 + ep_id*1_000_003 + query_idx), and
XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0.
Two settings that cost us hours
XLA_PYTHON_CLIENT_MEM_FRACTION=0.35, not 0.9. At 0.9 JAX takes 43 GB of a 48 GB L40S and the MuJoCo/EGL contexts die withGL_FRAMEBUFFER_UNSUPPORTED(0x8cdd) once you run more than ~2 workers.- mp4 round-trip belongs inside the
if not action_plan:branch. Outside it, the encode runs every step and is discarded on 4 of 5 β 3.8Γ slower, measured on identical episodes with identical outcomes.
Parallelism: one server per GPU, K=4 workers, sharded by --args.env_seed_offset.
Verified bit-identical noise and 100% outcome agreement vs serial β
verify_parallel_seeding.py. ~2.5Γ aggregate speedup, ~2.2 h per 160-episode target.
Reproducibility
Everything the code controls is exact. The residual is MuJoCo EGL rasterisation, which is not bit-reproducible across processes (~1 pixel / 1 LSB). Measured: 4/8 episodes bit-identical across cold processes, 8/8 agreeing on outcome, identical aggregate. Long rollouts diverge more often; one 740-step episode was bit-identical throughout.
Report success rates over the full set, never per-episode outcomes.
Comparability caveats
- MIMICGEN has 0 AluminumFoil006 training episodes (VACE has 29) and only 5 Jar025 (VACE 30). Those 40 episodes are effectively zero-shot for it. Use the common-140 subset as the head-to-head; report the 20 AluminumFoil006 episodes separately.
- POSE6DAUG stopped at 5000 steps; VACE/MIMICGEN ran to 30000. Compare only at
matched steps (1000/1500/5000), where all three saw the same LR schedule
(
decay_steps=30_000deliberately unchanged). - 20 episodes per object β one flip is 5 points. Per-object differences of one or two successes are noise.