transfer / eval_stack /README.md
Ronaldo-GOAT's picture
eval_stack: point REPLAY_ROOT at the in-repo negmesh8 bank
6f060de verified
|
Raw History Blame Contribute Delete
8.5 kB
# pi0.5 RoboCasa eval stack β€” negmesh8 160-episode replay benchmark
Everything needed to reproduce our evaluation on another cluster. Built and verified
2026-09-22/23 on lab-gpu26.
## What this evaluates
`PickPlaceCounterToCabinet`, 8 "hard mesh" objects Γ— 20 exact-replay episodes = **160**.
Three augmentation arms fine-tuned from the same `pi05_ctc502_60k` warm start, plus the
base itself as reference.
## Results so far (L40S, `per_query_epid`, action seed 42)
| checkpoint | all-160 |
|---|---|
| **base `pi05_ctc502_60k`** | **17/160 = 10.62%** |
| published reference for the same 160 | 7/160 = 4.38% β€” **not comparable, see below** |
Per object (base): AluminumFoil006 0/20 Β· BlenderJug023 3/20 Β· BlenderJug024 5/20 Β·
Jar025 4/20 Β· Juice008 0/20 Β· SyrupBottle006 2/20 Β· teapot_7 2/20 Β· wine_5 1/20
## THE THREE SERVING BUGS β€” read before running anything
The published 4.38% came from a harness that served the checkpoint wrongly. All three
must be right or your numbers are meaningless:
| | wrong (published) | correct |
|---|---|---|
| `action_horizon` | 10 | **50** |
| normalization | quantile (q01/q99) | **z-score** for the base |
| norm stats | `robocasa365_human300` | **the checkpoint's own ctc502 stats** |
Measured impact: **1/8 β†’ 5/8** on identical seeded episodes; **7/160 β†’ 17/160** on the
replay set. `jobs/robocasa_eval/assert_norm_config.py` gates every job against this β€”
run it, don't skip it.
**Fine-tuned arms differ from the base here**: VACE / MIMICGEN / POSE6DAUG were trained
with `use_quantile_norm=True` on `ctc502_qnorm`, so they must be served **quantile**,
not z-score. Each ships its own `assets/ctc502_qnorm/`. Crossing these silently corrupts
results.
## TWO different "VACE" checkpoint sets exist β€” do not cross them
| HF repo | run | init | data | ships `assets/` | serve with |
|---|---|---|---|---|---|
| `pi05-VACE-negmesh8-243ep-ctc502-60k-30k-h100x2` | 19500 | `pi05_ctc502_60k` | 243 eps (drop13) | `ctc502_qnorm` | `..._vace_negmesh` (quantile) |
| `pi05-transfer-vace-negmesh256-allintra-bs64-30k-h100x4` | 18808 | **`pi05_base`** | **256 eps** | **`robocasa365_human300`** | a config with `asset_id="robocasa365_human300"` |
Only the first is the VACE arm of the three-way comparison in this package. The second is an
earlier run on a different initialization, different data and different normalization; its
numbers are not comparable to anything here.
Each checkpoint ships its own `assets/<asset_id>/`, so `policy_config.py` resolves from the
checkpoint itself β€” serving one with the other's config **fails loudly** (`norm stats not
found ... checkpoint assets/ contains: [...]`) rather than corrupting silently. Run
`jobs/robocasa_eval/assert_norm_config.py` and trust it.
## GPU ARCHITECTURE IS A CONTROLLED VARIABLE
A100 vs L40S flips ~2 of 8 episodes. Every cross-architecture episode pair diverges at
**step 0, query 0** β€” the initial EGL rasterisation differs. Within one architecture,
two different physical GPUs agreed on 8/8 outcomes.
**Run every checkpoint you intend to compare on the same GPU model.** Our numbers are
L40S. `summarize_shards.py` refuses to aggregate across mixed `gpu_model` values.
## Setup
1. `openpi` from `Ronaldo-GOAT/transfer` (`transfer.tar`) β€” includes the **vendored
robosuite** with `load_model_on_init`. The PyPI robosuite 1.5.2 reports the same
version string but lacks it and will fail `gym.make`. Uninstall PyPI robosuite, then
`pip install -e third_party/robosuite`, and verify `robosuite.__file__` is inside the
bundle.
2. RoboCasa assets (~16 GB extracted, 123,586 files) β€” `jobs/robocasa_eval/fetch_robocasa_assets.py`.
Not bundled. robocasa's own `download_kitchen_assets.py` calls `input()` and omits the
fixtures archive; use ours.
3. **The 160-episode replay bank β€” get the right one.** It now ships **in this repo** at
`eval_stack/replay_bank_negmesh8/episodes/<object>/episode_0NN/` (968 files, 40.6 MB).
Mirror of `Ronaldo-GOAT/pi05-neg-mesh:episodes/`. 8 objects x 20 episodes, env seeds
10000-10019, **policy/action seed 42**. Each episode ships `initial_model.xml.gz`,
`initial_sim_state.npz`, `initial_observation.npz`, `episode_metadata.pkl`,
`replay_state.json`, `result.json`. ~40 MB total.
**Verify you have the right bank** β€” `ls episodes/` must be exactly these 8:
```
aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023
blender_jug__BlenderJug024 jar__Jar025
juice__Juice008 syrup_bottle__SyrupBottle006
teapot__teapot_7 wine__wine_5
```
ep_id -> object is contiguous blocks of 20 in that alphabetical order: 0-19 AluminumFoil006,
20-39 BlenderJug023, 40-59 BlenderJug024, 60-79 Jar025, 80-99 Juice008,
100-119 SyrupBottle006, 120-139 teapot_7, 140-159 wine_5.
**There is a second, incompatible 160-episode bank** shipped in the same upstream repo:
`actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_*`. Its objects are
`donut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6,
teapot_7` β€” **only `teapot_7` overlaps ours**, and its policy seed is **12345**, not 42.
It cannot score these checkpoints, and its `eval_exact160.sh` is a GR00T harness that
cannot even load them. Watch for the near-collisions: Jar**023**/Jar**025**,
SyrupBottle**005**/**006**, teapot_**6**/**7**. See `docs/EVAL_SETS.md`.
4. Overlay `harness/` onto the openpi tree:
`main.py`, `get_eval_stats.py` β†’ `examples/robocasa/` Β· `serve_policy.py` β†’ `scripts/` Β·
`policy.py`, `policy_config.py` β†’ `src/openpi/policies/` Β· `config.py` β†’ `src/openpi/training/`
## Document precedence
`README.md` (this file) > `docs/EVAL_SETS.md` > `docs/EVAL_SEEDS.md`. The last is the
upstream doc kept for provenance; it states `action_horizon=10`, which is **wrong** for
every checkpoint here. Its header now says so.
## Run β€” the 160-episode replay benchmark (what all our numbers use)
```bash
CKPT_DIR=<step_dir> # containing params/ and assets/
CONFIG_NAME=<see below>
REPLAY_ROOT=eval_stack/replay_bank_negmesh8/episodes
ENV_SEED_OFFSET=0 NUM_TRIALS=160
bash jobs/robocasa_eval.sbatch
python jobs/robocasa_eval/summarize_shards.py --dir <out>
```
Configs: base β†’ `pi05_robocasa_ctc502_60k_eval` (z-score) Β· VACE β†’
`..._vace_negmesh` Β· MIMICGEN β†’ `..._mimicgen_aug256` Β· POSE6DAUG β†’ `..._pose6daug_neg256`
(all quantile).
Fixed protocol: `action_seed=42`, `replan_steps=5`, `resize_size=224`, camera 256,
`mp4_roundtrip=True`, `generative_textures=False`, noise
`default_rng(42 + ep_id*1_000_003 + query_idx)`, and
`XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0`.
## Two settings that cost us hours
- **`XLA_PYTHON_CLIENT_MEM_FRACTION=0.35`**, not 0.9. At 0.9 JAX takes 43 GB of a 48 GB
L40S and the MuJoCo/EGL contexts die with `GL_FRAMEBUFFER_UNSUPPORTED` (0x8cdd) once
you run more than ~2 workers.
- **mp4 round-trip belongs inside the `if not action_plan:` branch.** Outside it, the
encode runs every step and is discarded on 4 of 5 β€” **3.8Γ— slower**, measured on
identical episodes with identical outcomes.
Parallelism: one server per GPU, K=4 workers, sharded by `--args.env_seed_offset`.
Verified bit-identical noise and 100% outcome agreement vs serial β€”
`verify_parallel_seeding.py`. ~2.5Γ— aggregate speedup, ~2.2 h per 160-episode target.
## Reproducibility
Everything the code controls is exact. The residual is MuJoCo EGL rasterisation, which is
not bit-reproducible across processes (~1 pixel / 1 LSB). Measured: 4/8 episodes
bit-identical across cold processes, **8/8 agreeing on outcome, identical aggregate**.
Long rollouts diverge more often; one 740-step episode was bit-identical throughout.
**Report success rates over the full set, never per-episode outcomes.**
## Comparability caveats
- **MIMICGEN has 0 AluminumFoil006 training episodes** (VACE has 29) and only 5 Jar025
(VACE 30). Those 40 episodes are effectively zero-shot for it. Use the **common-140**
subset as the head-to-head; report the 20 AluminumFoil006 episodes separately.
- **POSE6DAUG stopped at 5000 steps**; VACE/MIMICGEN ran to 30000. Compare only at
matched steps (1000/1500/5000), where all three saw the same LR schedule
(`decay_steps=30_000` deliberately unchanged).
- 20 episodes per object β‡’ one flip is 5 points. Per-object differences of one or two
successes are noise.