File size: 8,496 Bytes
51c2c72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c4f9c9a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
51c2c72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6f060de
 
 
9bfcc29
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
51c2c72
 
 
 
b963255
 
 
 
 
 
 
51c2c72
 
 
 
6f060de
51c2c72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
# pi0.5 RoboCasa eval stack β€” negmesh8 160-episode replay benchmark

Everything needed to reproduce our evaluation on another cluster. Built and verified
2026-09-22/23 on lab-gpu26.

## What this evaluates

`PickPlaceCounterToCabinet`, 8 "hard mesh" objects Γ— 20 exact-replay episodes = **160**.
Three augmentation arms fine-tuned from the same `pi05_ctc502_60k` warm start, plus the
base itself as reference.

## Results so far (L40S, `per_query_epid`, action seed 42)

| checkpoint | all-160 |
|---|---|
| **base `pi05_ctc502_60k`** | **17/160 = 10.62%** |
| published reference for the same 160 | 7/160 = 4.38% β€” **not comparable, see below** |

Per object (base): AluminumFoil006 0/20 Β· BlenderJug023 3/20 Β· BlenderJug024 5/20 Β·
Jar025 4/20 Β· Juice008 0/20 Β· SyrupBottle006 2/20 Β· teapot_7 2/20 Β· wine_5 1/20

## THE THREE SERVING BUGS β€” read before running anything

The published 4.38% came from a harness that served the checkpoint wrongly. All three
must be right or your numbers are meaningless:

| | wrong (published) | correct |
|---|---|---|
| `action_horizon` | 10 | **50** |
| normalization | quantile (q01/q99) | **z-score** for the base |
| norm stats | `robocasa365_human300` | **the checkpoint's own ctc502 stats** |

Measured impact: **1/8 β†’ 5/8** on identical seeded episodes; **7/160 β†’ 17/160** on the
replay set. `jobs/robocasa_eval/assert_norm_config.py` gates every job against this β€”
run it, don't skip it.

**Fine-tuned arms differ from the base here**: VACE / MIMICGEN / POSE6DAUG were trained
with `use_quantile_norm=True` on `ctc502_qnorm`, so they must be served **quantile**,
not z-score. Each ships its own `assets/ctc502_qnorm/`. Crossing these silently corrupts
results.


## TWO different "VACE" checkpoint sets exist β€” do not cross them

| HF repo | run | init | data | ships `assets/` | serve with |
|---|---|---|---|---|---|
| `pi05-VACE-negmesh8-243ep-ctc502-60k-30k-h100x2` | 19500 | `pi05_ctc502_60k` | 243 eps (drop13) | `ctc502_qnorm` | `..._vace_negmesh` (quantile) |
| `pi05-transfer-vace-negmesh256-allintra-bs64-30k-h100x4` | 18808 | **`pi05_base`** | **256 eps** | **`robocasa365_human300`** | a config with `asset_id="robocasa365_human300"` |

Only the first is the VACE arm of the three-way comparison in this package. The second is an
earlier run on a different initialization, different data and different normalization; its
numbers are not comparable to anything here.

Each checkpoint ships its own `assets/<asset_id>/`, so `policy_config.py` resolves from the
checkpoint itself β€” serving one with the other's config **fails loudly** (`norm stats not
found ... checkpoint assets/ contains: [...]`) rather than corrupting silently. Run
`jobs/robocasa_eval/assert_norm_config.py` and trust it.

## GPU ARCHITECTURE IS A CONTROLLED VARIABLE

A100 vs L40S flips ~2 of 8 episodes. Every cross-architecture episode pair diverges at
**step 0, query 0** β€” the initial EGL rasterisation differs. Within one architecture,
two different physical GPUs agreed on 8/8 outcomes.

**Run every checkpoint you intend to compare on the same GPU model.** Our numbers are
L40S. `summarize_shards.py` refuses to aggregate across mixed `gpu_model` values.

## Setup

1. `openpi` from `Ronaldo-GOAT/transfer` (`transfer.tar`) β€” includes the **vendored
   robosuite** with `load_model_on_init`. The PyPI robosuite 1.5.2 reports the same
   version string but lacks it and will fail `gym.make`. Uninstall PyPI robosuite, then
   `pip install -e third_party/robosuite`, and verify `robosuite.__file__` is inside the
   bundle.
2. RoboCasa assets (~16 GB extracted, 123,586 files) β€” `jobs/robocasa_eval/fetch_robocasa_assets.py`.
   Not bundled. robocasa's own `download_kitchen_assets.py` calls `input()` and omits the
   fixtures archive; use ours.
3. **The 160-episode replay bank β€” get the right one.** It now ships **in this repo** at
   `eval_stack/replay_bank_negmesh8/episodes/<object>/episode_0NN/` (968 files, 40.6 MB).
   Mirror of `Ronaldo-GOAT/pi05-neg-mesh:episodes/`. 8 objects x 20 episodes, env seeds
   10000-10019, **policy/action seed 42**. Each episode ships `initial_model.xml.gz`,
   `initial_sim_state.npz`, `initial_observation.npz`, `episode_metadata.pkl`,
   `replay_state.json`, `result.json`. ~40 MB total.

   **Verify you have the right bank** β€” `ls episodes/` must be exactly these 8:
   ```
   aluminum_foil__AluminumFoil006   blender_jug__BlenderJug023
   blender_jug__BlenderJug024       jar__Jar025
   juice__Juice008                  syrup_bottle__SyrupBottle006
   teapot__teapot_7                 wine__wine_5
   ```
   ep_id -> object is contiguous blocks of 20 in that alphabetical order: 0-19 AluminumFoil006,
   20-39 BlenderJug023, 40-59 BlenderJug024, 60-79 Jar025, 80-99 Juice008,
   100-119 SyrupBottle006, 120-139 teapot_7, 140-159 wine_5.

   **There is a second, incompatible 160-episode bank** shipped in the same upstream repo:
   `actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_*`. Its objects are
   `donut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6,
   teapot_7` β€” **only `teapot_7` overlaps ours**, and its policy seed is **12345**, not 42.
   It cannot score these checkpoints, and its `eval_exact160.sh` is a GR00T harness that
   cannot even load them. Watch for the near-collisions: Jar**023**/Jar**025**,
   SyrupBottle**005**/**006**, teapot_**6**/**7**. See `docs/EVAL_SETS.md`.
4. Overlay `harness/` onto the openpi tree:
   `main.py`, `get_eval_stats.py` β†’ `examples/robocasa/` Β· `serve_policy.py` β†’ `scripts/` Β·
   `policy.py`, `policy_config.py` β†’ `src/openpi/policies/` Β· `config.py` β†’ `src/openpi/training/`

## Document precedence

`README.md` (this file) > `docs/EVAL_SETS.md` > `docs/EVAL_SEEDS.md`. The last is the
upstream doc kept for provenance; it states `action_horizon=10`, which is **wrong** for
every checkpoint here. Its header now says so.

## Run β€” the 160-episode replay benchmark (what all our numbers use)

```bash
CKPT_DIR=<step_dir>            # containing params/ and assets/
CONFIG_NAME=<see below>
REPLAY_ROOT=eval_stack/replay_bank_negmesh8/episodes
ENV_SEED_OFFSET=0 NUM_TRIALS=160
bash jobs/robocasa_eval.sbatch
python jobs/robocasa_eval/summarize_shards.py --dir <out>
```

Configs: base β†’ `pi05_robocasa_ctc502_60k_eval` (z-score) Β· VACE β†’
`..._vace_negmesh` Β· MIMICGEN β†’ `..._mimicgen_aug256` Β· POSE6DAUG β†’ `..._pose6daug_neg256`
(all quantile).

Fixed protocol: `action_seed=42`, `replan_steps=5`, `resize_size=224`, camera 256,
`mp4_roundtrip=True`, `generative_textures=False`, noise
`default_rng(42 + ep_id*1_000_003 + query_idx)`, and
`XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0`.

## Two settings that cost us hours

- **`XLA_PYTHON_CLIENT_MEM_FRACTION=0.35`**, not 0.9. At 0.9 JAX takes 43 GB of a 48 GB
  L40S and the MuJoCo/EGL contexts die with `GL_FRAMEBUFFER_UNSUPPORTED` (0x8cdd) once
  you run more than ~2 workers.
- **mp4 round-trip belongs inside the `if not action_plan:` branch.** Outside it, the
  encode runs every step and is discarded on 4 of 5 β€” **3.8Γ— slower**, measured on
  identical episodes with identical outcomes.

Parallelism: one server per GPU, K=4 workers, sharded by `--args.env_seed_offset`.
Verified bit-identical noise and 100% outcome agreement vs serial β€”
`verify_parallel_seeding.py`. ~2.5Γ— aggregate speedup, ~2.2 h per 160-episode target.

## Reproducibility

Everything the code controls is exact. The residual is MuJoCo EGL rasterisation, which is
not bit-reproducible across processes (~1 pixel / 1 LSB). Measured: 4/8 episodes
bit-identical across cold processes, **8/8 agreeing on outcome, identical aggregate**.
Long rollouts diverge more often; one 740-step episode was bit-identical throughout.

**Report success rates over the full set, never per-episode outcomes.**

## Comparability caveats

- **MIMICGEN has 0 AluminumFoil006 training episodes** (VACE has 29) and only 5 Jar025
  (VACE 30). Those 40 episodes are effectively zero-shot for it. Use the **common-140**
  subset as the head-to-head; report the 20 AluminumFoil006 episodes separately.
- **POSE6DAUG stopped at 5000 steps**; VACE/MIMICGEN ran to 30000. Compare only at
  matched steps (1000/1500/5000), where all three saw the same LR schedule
  (`decay_steps=30_000` deliberately unchanged).
- 20 episodes per object β‡’ one flip is 5 points. Per-object differences of one or two
  successes are noise.