|
Download eval_stack/README.md from Ronaldo-GOAT/transfer: direct link, hf CLI and curl.
- Browser
- Download file 8.5 kB
-
https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/README.md
- Command line
-
hf download hf://Ronaldo-GOAT/transfer/eval_stack/README.md
-
curl -L -o README.md https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/README.md
8.5 kB
| # pi0.5 RoboCasa eval stack β negmesh8 160-episode replay benchmark | |
| Everything needed to reproduce our evaluation on another cluster. Built and verified | |
| 2026-09-22/23 on lab-gpu26. | |
| ## What this evaluates | |
| `PickPlaceCounterToCabinet`, 8 "hard mesh" objects Γ 20 exact-replay episodes = **160**. | |
| Three augmentation arms fine-tuned from the same `pi05_ctc502_60k` warm start, plus the | |
| base itself as reference. | |
| ## Results so far (L40S, `per_query_epid`, action seed 42) | |
| | checkpoint | all-160 | | |
| |---|---| | |
| | **base `pi05_ctc502_60k`** | **17/160 = 10.62%** | | |
| | published reference for the same 160 | 7/160 = 4.38% β **not comparable, see below** | | |
| Per object (base): AluminumFoil006 0/20 Β· BlenderJug023 3/20 Β· BlenderJug024 5/20 Β· | |
| Jar025 4/20 Β· Juice008 0/20 Β· SyrupBottle006 2/20 Β· teapot_7 2/20 Β· wine_5 1/20 | |
| ## THE THREE SERVING BUGS β read before running anything | |
| The published 4.38% came from a harness that served the checkpoint wrongly. All three | |
| must be right or your numbers are meaningless: | |
| | | wrong (published) | correct | | |
| |---|---|---| | |
| | `action_horizon` | 10 | **50** | | |
| | normalization | quantile (q01/q99) | **z-score** for the base | | |
| | norm stats | `robocasa365_human300` | **the checkpoint's own ctc502 stats** | | |
| Measured impact: **1/8 β 5/8** on identical seeded episodes; **7/160 β 17/160** on the | |
| replay set. `jobs/robocasa_eval/assert_norm_config.py` gates every job against this β | |
| run it, don't skip it. | |
| **Fine-tuned arms differ from the base here**: VACE / MIMICGEN / POSE6DAUG were trained | |
| with `use_quantile_norm=True` on `ctc502_qnorm`, so they must be served **quantile**, | |
| not z-score. Each ships its own `assets/ctc502_qnorm/`. Crossing these silently corrupts | |
| results. | |
| ## TWO different "VACE" checkpoint sets exist β do not cross them | |
| | HF repo | run | init | data | ships `assets/` | serve with | | |
| |---|---|---|---|---|---| | |
| | `pi05-VACE-negmesh8-243ep-ctc502-60k-30k-h100x2` | 19500 | `pi05_ctc502_60k` | 243 eps (drop13) | `ctc502_qnorm` | `..._vace_negmesh` (quantile) | | |
| | `pi05-transfer-vace-negmesh256-allintra-bs64-30k-h100x4` | 18808 | **`pi05_base`** | **256 eps** | **`robocasa365_human300`** | a config with `asset_id="robocasa365_human300"` | | |
| Only the first is the VACE arm of the three-way comparison in this package. The second is an | |
| earlier run on a different initialization, different data and different normalization; its | |
| numbers are not comparable to anything here. | |
| Each checkpoint ships its own `assets/<asset_id>/`, so `policy_config.py` resolves from the | |
| checkpoint itself β serving one with the other's config **fails loudly** (`norm stats not | |
| found ... checkpoint assets/ contains: [...]`) rather than corrupting silently. Run | |
| `jobs/robocasa_eval/assert_norm_config.py` and trust it. | |
| ## GPU ARCHITECTURE IS A CONTROLLED VARIABLE | |
| A100 vs L40S flips ~2 of 8 episodes. Every cross-architecture episode pair diverges at | |
| **step 0, query 0** β the initial EGL rasterisation differs. Within one architecture, | |
| two different physical GPUs agreed on 8/8 outcomes. | |
| **Run every checkpoint you intend to compare on the same GPU model.** Our numbers are | |
| L40S. `summarize_shards.py` refuses to aggregate across mixed `gpu_model` values. | |
| ## Setup | |
| 1. `openpi` from `Ronaldo-GOAT/transfer` (`transfer.tar`) β includes the **vendored | |
| robosuite** with `load_model_on_init`. The PyPI robosuite 1.5.2 reports the same | |
| version string but lacks it and will fail `gym.make`. Uninstall PyPI robosuite, then | |
| `pip install -e third_party/robosuite`, and verify `robosuite.__file__` is inside the | |
| bundle. | |
| 2. RoboCasa assets (~16 GB extracted, 123,586 files) β `jobs/robocasa_eval/fetch_robocasa_assets.py`. | |
| Not bundled. robocasa's own `download_kitchen_assets.py` calls `input()` and omits the | |
| fixtures archive; use ours. | |
| 3. **The 160-episode replay bank β get the right one.** It now ships **in this repo** at | |
| `eval_stack/replay_bank_negmesh8/episodes/<object>/episode_0NN/` (968 files, 40.6 MB). | |
| Mirror of `Ronaldo-GOAT/pi05-neg-mesh:episodes/`. 8 objects x 20 episodes, env seeds | |
| 10000-10019, **policy/action seed 42**. Each episode ships `initial_model.xml.gz`, | |
| `initial_sim_state.npz`, `initial_observation.npz`, `episode_metadata.pkl`, | |
| `replay_state.json`, `result.json`. ~40 MB total. | |
| **Verify you have the right bank** β `ls episodes/` must be exactly these 8: | |
| ``` | |
| aluminum_foil__AluminumFoil006 blender_jug__BlenderJug023 | |
| blender_jug__BlenderJug024 jar__Jar025 | |
| juice__Juice008 syrup_bottle__SyrupBottle006 | |
| teapot__teapot_7 wine__wine_5 | |
| ``` | |
| ep_id -> object is contiguous blocks of 20 in that alphabetical order: 0-19 AluminumFoil006, | |
| 20-39 BlenderJug023, 40-59 BlenderJug024, 60-79 Jar025, 80-99 Juice008, | |
| 100-119 SyrupBottle006, 120-139 teapot_7, 140-159 wine_5. | |
| **There is a second, incompatible 160-episode bank** shipped in the same upstream repo: | |
| `actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_*`. Its objects are | |
| `donut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6, | |
| teapot_7` β **only `teapot_7` overlaps ours**, and its policy seed is **12345**, not 42. | |
| It cannot score these checkpoints, and its `eval_exact160.sh` is a GR00T harness that | |
| cannot even load them. Watch for the near-collisions: Jar**023**/Jar**025**, | |
| SyrupBottle**005**/**006**, teapot_**6**/**7**. See `docs/EVAL_SETS.md`. | |
| 4. Overlay `harness/` onto the openpi tree: | |
| `main.py`, `get_eval_stats.py` β `examples/robocasa/` Β· `serve_policy.py` β `scripts/` Β· | |
| `policy.py`, `policy_config.py` β `src/openpi/policies/` Β· `config.py` β `src/openpi/training/` | |
| ## Document precedence | |
| `README.md` (this file) > `docs/EVAL_SETS.md` > `docs/EVAL_SEEDS.md`. The last is the | |
| upstream doc kept for provenance; it states `action_horizon=10`, which is **wrong** for | |
| every checkpoint here. Its header now says so. | |
| ## Run β the 160-episode replay benchmark (what all our numbers use) | |
| ```bash | |
| CKPT_DIR=<step_dir> # containing params/ and assets/ | |
| CONFIG_NAME=<see below> | |
| REPLAY_ROOT=eval_stack/replay_bank_negmesh8/episodes | |
| ENV_SEED_OFFSET=0 NUM_TRIALS=160 | |
| bash jobs/robocasa_eval.sbatch | |
| python jobs/robocasa_eval/summarize_shards.py --dir <out> | |
| ``` | |
| Configs: base β `pi05_robocasa_ctc502_60k_eval` (z-score) Β· VACE β | |
| `..._vace_negmesh` Β· MIMICGEN β `..._mimicgen_aug256` Β· POSE6DAUG β `..._pose6daug_neg256` | |
| (all quantile). | |
| Fixed protocol: `action_seed=42`, `replan_steps=5`, `resize_size=224`, camera 256, | |
| `mp4_roundtrip=True`, `generative_textures=False`, noise | |
| `default_rng(42 + ep_id*1_000_003 + query_idx)`, and | |
| `XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0`. | |
| ## Two settings that cost us hours | |
| - **`XLA_PYTHON_CLIENT_MEM_FRACTION=0.35`**, not 0.9. At 0.9 JAX takes 43 GB of a 48 GB | |
| L40S and the MuJoCo/EGL contexts die with `GL_FRAMEBUFFER_UNSUPPORTED` (0x8cdd) once | |
| you run more than ~2 workers. | |
| - **mp4 round-trip belongs inside the `if not action_plan:` branch.** Outside it, the | |
| encode runs every step and is discarded on 4 of 5 β **3.8Γ slower**, measured on | |
| identical episodes with identical outcomes. | |
| Parallelism: one server per GPU, K=4 workers, sharded by `--args.env_seed_offset`. | |
| Verified bit-identical noise and 100% outcome agreement vs serial β | |
| `verify_parallel_seeding.py`. ~2.5Γ aggregate speedup, ~2.2 h per 160-episode target. | |
| ## Reproducibility | |
| Everything the code controls is exact. The residual is MuJoCo EGL rasterisation, which is | |
| not bit-reproducible across processes (~1 pixel / 1 LSB). Measured: 4/8 episodes | |
| bit-identical across cold processes, **8/8 agreeing on outcome, identical aggregate**. | |
| Long rollouts diverge more often; one 740-step episode was bit-identical throughout. | |
| **Report success rates over the full set, never per-episode outcomes.** | |
| ## Comparability caveats | |
| - **MIMICGEN has 0 AluminumFoil006 training episodes** (VACE has 29) and only 5 Jar025 | |
| (VACE 30). Those 40 episodes are effectively zero-shot for it. Use the **common-140** | |
| subset as the head-to-head; report the 20 AluminumFoil006 episodes separately. | |
| - **POSE6DAUG stopped at 5000 steps**; VACE/MIMICGEN ran to 30000. Compare only at | |
| matched steps (1000/1500/5000), where all three saw the same LR schedule | |
| (`decay_steps=30_000` deliberately unchanged). | |
| - 20 episodes per object β one flip is 5 points. Per-object differences of one or two | |
| successes are noise. | |