# pi0.5 RoboCasa eval stack — negmesh8 160-episode replay benchmark Everything needed to reproduce our evaluation on another cluster. Built and verified 2026-09-22/23 on lab-gpu26. ## What this evaluates `PickPlaceCounterToCabinet`, 8 "hard mesh" objects × 20 exact-replay episodes = **160**. Three augmentation arms fine-tuned from the same `pi05_ctc502_60k` warm start, plus the base itself as reference. ## Results so far (L40S, `per_query_epid`, action seed 42) | checkpoint | all-160 | |---|---| | **base `pi05_ctc502_60k`** | **17/160 = 10.62%** | | published reference for the same 160 | 7/160 = 4.38% — **not comparable, see below** | Per object (base): AluminumFoil006 0/20 · BlenderJug023 3/20 · BlenderJug024 5/20 · Jar025 4/20 · Juice008 0/20 · SyrupBottle006 2/20 · teapot_7 2/20 · wine_5 1/20 ## THE THREE SERVING BUGS — read before running anything The published 4.38% came from a harness that served the checkpoint wrongly. All three must be right or your numbers are meaningless: | | wrong (published) | correct | |---|---|---| | `action_horizon` | 10 | **50** | | normalization | quantile (q01/q99) | **z-score** for the base | | norm stats | `robocasa365_human300` | **the checkpoint's own ctc502 stats** | Measured impact: **1/8 → 5/8** on identical seeded episodes; **7/160 → 17/160** on the replay set. `jobs/robocasa_eval/assert_norm_config.py` gates every job against this — run it, don't skip it. **Fine-tuned arms differ from the base here**: VACE / MIMICGEN / POSE6DAUG were trained with `use_quantile_norm=True` on `ctc502_qnorm`, so they must be served **quantile**, not z-score. Each ships its own `assets/ctc502_qnorm/`. Crossing these silently corrupts results. ## TWO different "VACE" checkpoint sets exist — do not cross them | HF repo | run | init | data | ships `assets/` | serve with | |---|---|---|---|---|---| | `pi05-VACE-negmesh8-243ep-ctc502-60k-30k-h100x2` | 19500 | `pi05_ctc502_60k` | 243 eps (drop13) | `ctc502_qnorm` | `..._vace_negmesh` (quantile) | | `pi05-transfer-vace-negmesh256-allintra-bs64-30k-h100x4` | 18808 | **`pi05_base`** | **256 eps** | **`robocasa365_human300`** | a config with `asset_id="robocasa365_human300"` | Only the first is the VACE arm of the three-way comparison in this package. The second is an earlier run on a different initialization, different data and different normalization; its numbers are not comparable to anything here. Each checkpoint ships its own `assets//`, so `policy_config.py` resolves from the checkpoint itself — serving one with the other's config **fails loudly** (`norm stats not found ... checkpoint assets/ contains: [...]`) rather than corrupting silently. Run `jobs/robocasa_eval/assert_norm_config.py` and trust it. ## GPU ARCHITECTURE IS A CONTROLLED VARIABLE A100 vs L40S flips ~2 of 8 episodes. Every cross-architecture episode pair diverges at **step 0, query 0** — the initial EGL rasterisation differs. Within one architecture, two different physical GPUs agreed on 8/8 outcomes. **Run every checkpoint you intend to compare on the same GPU model.** Our numbers are L40S. `summarize_shards.py` refuses to aggregate across mixed `gpu_model` values. ## Setup 1. `openpi` from `Ronaldo-GOAT/transfer` (`transfer.tar`) — includes the **vendored robosuite** with `load_model_on_init`. The PyPI robosuite 1.5.2 reports the same version string but lacks it and will fail `gym.make`. Uninstall PyPI robosuite, then `pip install -e third_party/robosuite`, and verify `robosuite.__file__` is inside the bundle. 2. RoboCasa assets (~16 GB extracted, 123,586 files) — `jobs/robocasa_eval/fetch_robocasa_assets.py`. Not bundled. robocasa's own `download_kitchen_assets.py` calls `input()` and omits the fixtures archive; use ours. 3. **The 160-episode replay bank — get the right one.** It now ships **in this repo** at `eval_stack/replay_bank_negmesh8/episodes//episode_0NN/` (968 files, 40.6 MB). Mirror of `Ronaldo-GOAT/pi05-neg-mesh:episodes/`. 8 objects x 20 episodes, env seeds 10000-10019, **policy/action seed 42**. Each episode ships `initial_model.xml.gz`, `initial_sim_state.npz`, `initial_observation.npz`, `episode_metadata.pkl`, `replay_state.json`, `result.json`. ~40 MB total. **Verify you have the right bank** — `ls episodes/` must be exactly these 8: ``` aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023 blender_jug__BlenderJug024 jar__Jar025 juice__Juice008 syrup_bottle__SyrupBottle006 teapot__teapot_7 wine__wine_5 ``` ep_id -> object is contiguous blocks of 20 in that alphabetical order: 0-19 AluminumFoil006, 20-39 BlenderJug023, 40-59 BlenderJug024, 60-79 Jar025, 80-99 Juice008, 100-119 SyrupBottle006, 120-139 teapot_7, 140-159 wine_5. **There is a second, incompatible 160-episode bank** shipped in the same upstream repo: `actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_*`. Its objects are `donut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6, teapot_7` — **only `teapot_7` overlaps ours**, and its policy seed is **12345**, not 42. It cannot score these checkpoints, and its `eval_exact160.sh` is a GR00T harness that cannot even load them. Watch for the near-collisions: Jar**023**/Jar**025**, SyrupBottle**005**/**006**, teapot_**6**/**7**. See `docs/EVAL_SETS.md`. 4. Overlay `harness/` onto the openpi tree: `main.py`, `get_eval_stats.py` → `examples/robocasa/` · `serve_policy.py` → `scripts/` · `policy.py`, `policy_config.py` → `src/openpi/policies/` · `config.py` → `src/openpi/training/` ## Document precedence `README.md` (this file) > `docs/EVAL_SETS.md` > `docs/EVAL_SEEDS.md`. The last is the upstream doc kept for provenance; it states `action_horizon=10`, which is **wrong** for every checkpoint here. Its header now says so. ## Run — the 160-episode replay benchmark (what all our numbers use) ```bash CKPT_DIR= # containing params/ and assets/ CONFIG_NAME= REPLAY_ROOT=eval_stack/replay_bank_negmesh8/episodes ENV_SEED_OFFSET=0 NUM_TRIALS=160 bash jobs/robocasa_eval.sbatch python jobs/robocasa_eval/summarize_shards.py --dir ``` Configs: base → `pi05_robocasa_ctc502_60k_eval` (z-score) · VACE → `..._vace_negmesh` · MIMICGEN → `..._mimicgen_aug256` · POSE6DAUG → `..._pose6daug_neg256` (all quantile). Fixed protocol: `action_seed=42`, `replan_steps=5`, `resize_size=224`, camera 256, `mp4_roundtrip=True`, `generative_textures=False`, noise `default_rng(42 + ep_id*1_000_003 + query_idx)`, and `XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0`. ## Two settings that cost us hours - **`XLA_PYTHON_CLIENT_MEM_FRACTION=0.35`**, not 0.9. At 0.9 JAX takes 43 GB of a 48 GB L40S and the MuJoCo/EGL contexts die with `GL_FRAMEBUFFER_UNSUPPORTED` (0x8cdd) once you run more than ~2 workers. - **mp4 round-trip belongs inside the `if not action_plan:` branch.** Outside it, the encode runs every step and is discarded on 4 of 5 — **3.8× slower**, measured on identical episodes with identical outcomes. Parallelism: one server per GPU, K=4 workers, sharded by `--args.env_seed_offset`. Verified bit-identical noise and 100% outcome agreement vs serial — `verify_parallel_seeding.py`. ~2.5× aggregate speedup, ~2.2 h per 160-episode target. ## Reproducibility Everything the code controls is exact. The residual is MuJoCo EGL rasterisation, which is not bit-reproducible across processes (~1 pixel / 1 LSB). Measured: 4/8 episodes bit-identical across cold processes, **8/8 agreeing on outcome, identical aggregate**. Long rollouts diverge more often; one 740-step episode was bit-identical throughout. **Report success rates over the full set, never per-episode outcomes.** ## Comparability caveats - **MIMICGEN has 0 AluminumFoil006 training episodes** (VACE has 29) and only 5 Jar025 (VACE 30). Those 40 episodes are effectively zero-shot for it. Use the **common-140** subset as the head-to-head; report the 20 AluminumFoil006 episodes separately. - **POSE6DAUG stopped at 5000 steps**; VACE/MIMICGEN ran to 30000. Compare only at matched steps (1000/1500/5000), where all three saw the same LR schedule (`decay_steps=30_000` deliberately unchanged). - 20 episodes per object ⇒ one flip is 5 points. Per-object differences of one or two successes are noise.