File size: 8,496 Bytes
51c2c72 c4f9c9a 51c2c72 6f060de 9bfcc29 51c2c72 b963255 51c2c72 6f060de 51c2c72 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 | # pi0.5 RoboCasa eval stack β negmesh8 160-episode replay benchmark
Everything needed to reproduce our evaluation on another cluster. Built and verified
2026-09-22/23 on lab-gpu26.
## What this evaluates
`PickPlaceCounterToCabinet`, 8 "hard mesh" objects Γ 20 exact-replay episodes = **160**.
Three augmentation arms fine-tuned from the same `pi05_ctc502_60k` warm start, plus the
base itself as reference.
## Results so far (L40S, `per_query_epid`, action seed 42)
| checkpoint | all-160 |
|---|---|
| **base `pi05_ctc502_60k`** | **17/160 = 10.62%** |
| published reference for the same 160 | 7/160 = 4.38% β **not comparable, see below** |
Per object (base): AluminumFoil006 0/20 Β· BlenderJug023 3/20 Β· BlenderJug024 5/20 Β·
Jar025 4/20 Β· Juice008 0/20 Β· SyrupBottle006 2/20 Β· teapot_7 2/20 Β· wine_5 1/20
## THE THREE SERVING BUGS β read before running anything
The published 4.38% came from a harness that served the checkpoint wrongly. All three
must be right or your numbers are meaningless:
| | wrong (published) | correct |
|---|---|---|
| `action_horizon` | 10 | **50** |
| normalization | quantile (q01/q99) | **z-score** for the base |
| norm stats | `robocasa365_human300` | **the checkpoint's own ctc502 stats** |
Measured impact: **1/8 β 5/8** on identical seeded episodes; **7/160 β 17/160** on the
replay set. `jobs/robocasa_eval/assert_norm_config.py` gates every job against this β
run it, don't skip it.
**Fine-tuned arms differ from the base here**: VACE / MIMICGEN / POSE6DAUG were trained
with `use_quantile_norm=True` on `ctc502_qnorm`, so they must be served **quantile**,
not z-score. Each ships its own `assets/ctc502_qnorm/`. Crossing these silently corrupts
results.
## TWO different "VACE" checkpoint sets exist β do not cross them
| HF repo | run | init | data | ships `assets/` | serve with |
|---|---|---|---|---|---|
| `pi05-VACE-negmesh8-243ep-ctc502-60k-30k-h100x2` | 19500 | `pi05_ctc502_60k` | 243 eps (drop13) | `ctc502_qnorm` | `..._vace_negmesh` (quantile) |
| `pi05-transfer-vace-negmesh256-allintra-bs64-30k-h100x4` | 18808 | **`pi05_base`** | **256 eps** | **`robocasa365_human300`** | a config with `asset_id="robocasa365_human300"` |
Only the first is the VACE arm of the three-way comparison in this package. The second is an
earlier run on a different initialization, different data and different normalization; its
numbers are not comparable to anything here.
Each checkpoint ships its own `assets/<asset_id>/`, so `policy_config.py` resolves from the
checkpoint itself β serving one with the other's config **fails loudly** (`norm stats not
found ... checkpoint assets/ contains: [...]`) rather than corrupting silently. Run
`jobs/robocasa_eval/assert_norm_config.py` and trust it.
## GPU ARCHITECTURE IS A CONTROLLED VARIABLE
A100 vs L40S flips ~2 of 8 episodes. Every cross-architecture episode pair diverges at
**step 0, query 0** β the initial EGL rasterisation differs. Within one architecture,
two different physical GPUs agreed on 8/8 outcomes.
**Run every checkpoint you intend to compare on the same GPU model.** Our numbers are
L40S. `summarize_shards.py` refuses to aggregate across mixed `gpu_model` values.
## Setup
1. `openpi` from `Ronaldo-GOAT/transfer` (`transfer.tar`) β includes the **vendored
robosuite** with `load_model_on_init`. The PyPI robosuite 1.5.2 reports the same
version string but lacks it and will fail `gym.make`. Uninstall PyPI robosuite, then
`pip install -e third_party/robosuite`, and verify `robosuite.__file__` is inside the
bundle.
2. RoboCasa assets (~16 GB extracted, 123,586 files) β `jobs/robocasa_eval/fetch_robocasa_assets.py`.
Not bundled. robocasa's own `download_kitchen_assets.py` calls `input()` and omits the
fixtures archive; use ours.
3. **The 160-episode replay bank β get the right one.** It now ships **in this repo** at
`eval_stack/replay_bank_negmesh8/episodes/<object>/episode_0NN/` (968 files, 40.6 MB).
Mirror of `Ronaldo-GOAT/pi05-neg-mesh:episodes/`. 8 objects x 20 episodes, env seeds
10000-10019, **policy/action seed 42**. Each episode ships `initial_model.xml.gz`,
`initial_sim_state.npz`, `initial_observation.npz`, `episode_metadata.pkl`,
`replay_state.json`, `result.json`. ~40 MB total.
**Verify you have the right bank** β `ls episodes/` must be exactly these 8:
```
aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023
blender_jug__BlenderJug024 jar__Jar025
juice__Juice008 syrup_bottle__SyrupBottle006
teapot__teapot_7 wine__wine_5
```
ep_id -> object is contiguous blocks of 20 in that alphabetical order: 0-19 AluminumFoil006,
20-39 BlenderJug023, 40-59 BlenderJug024, 60-79 Jar025, 80-99 Juice008,
100-119 SyrupBottle006, 120-139 teapot_7, 140-159 wine_5.
**There is a second, incompatible 160-episode bank** shipped in the same upstream repo:
`actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_*`. Its objects are
`donut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6,
teapot_7` β **only `teapot_7` overlaps ours**, and its policy seed is **12345**, not 42.
It cannot score these checkpoints, and its `eval_exact160.sh` is a GR00T harness that
cannot even load them. Watch for the near-collisions: Jar**023**/Jar**025**,
SyrupBottle**005**/**006**, teapot_**6**/**7**. See `docs/EVAL_SETS.md`.
4. Overlay `harness/` onto the openpi tree:
`main.py`, `get_eval_stats.py` β `examples/robocasa/` Β· `serve_policy.py` β `scripts/` Β·
`policy.py`, `policy_config.py` β `src/openpi/policies/` Β· `config.py` β `src/openpi/training/`
## Document precedence
`README.md` (this file) > `docs/EVAL_SETS.md` > `docs/EVAL_SEEDS.md`. The last is the
upstream doc kept for provenance; it states `action_horizon=10`, which is **wrong** for
every checkpoint here. Its header now says so.
## Run β the 160-episode replay benchmark (what all our numbers use)
```bash
CKPT_DIR=<step_dir> # containing params/ and assets/
CONFIG_NAME=<see below>
REPLAY_ROOT=eval_stack/replay_bank_negmesh8/episodes
ENV_SEED_OFFSET=0 NUM_TRIALS=160
bash jobs/robocasa_eval.sbatch
python jobs/robocasa_eval/summarize_shards.py --dir <out>
```
Configs: base β `pi05_robocasa_ctc502_60k_eval` (z-score) Β· VACE β
`..._vace_negmesh` Β· MIMICGEN β `..._mimicgen_aug256` Β· POSE6DAUG β `..._pose6daug_neg256`
(all quantile).
Fixed protocol: `action_seed=42`, `replan_steps=5`, `resize_size=224`, camera 256,
`mp4_roundtrip=True`, `generative_textures=False`, noise
`default_rng(42 + ep_id*1_000_003 + query_idx)`, and
`XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0`.
## Two settings that cost us hours
- **`XLA_PYTHON_CLIENT_MEM_FRACTION=0.35`**, not 0.9. At 0.9 JAX takes 43 GB of a 48 GB
L40S and the MuJoCo/EGL contexts die with `GL_FRAMEBUFFER_UNSUPPORTED` (0x8cdd) once
you run more than ~2 workers.
- **mp4 round-trip belongs inside the `if not action_plan:` branch.** Outside it, the
encode runs every step and is discarded on 4 of 5 β **3.8Γ slower**, measured on
identical episodes with identical outcomes.
Parallelism: one server per GPU, K=4 workers, sharded by `--args.env_seed_offset`.
Verified bit-identical noise and 100% outcome agreement vs serial β
`verify_parallel_seeding.py`. ~2.5Γ aggregate speedup, ~2.2 h per 160-episode target.
## Reproducibility
Everything the code controls is exact. The residual is MuJoCo EGL rasterisation, which is
not bit-reproducible across processes (~1 pixel / 1 LSB). Measured: 4/8 episodes
bit-identical across cold processes, **8/8 agreeing on outcome, identical aggregate**.
Long rollouts diverge more often; one 740-step episode was bit-identical throughout.
**Report success rates over the full set, never per-episode outcomes.**
## Comparability caveats
- **MIMICGEN has 0 AluminumFoil006 training episodes** (VACE has 29) and only 5 Jar025
(VACE 30). Those 40 episodes are effectively zero-shot for it. Use the **common-140**
subset as the head-to-head; report the 20 AluminumFoil006 episodes separately.
- **POSE6DAUG stopped at 5000 steps**; VACE/MIMICGEN ran to 30000. Compare only at
matched steps (1000/1500/5000), where all three saw the same LR schedule
(`decay_steps=30_000` deliberately unchanged).
- 20 episodes per object β one flip is 5 points. Per-object differences of one or two
successes are noise.
|