File size: 9,772 Bytes
b963255
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
51c2c72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
> **SUPERSEDED IN PART — read `../README.md` first.**
>
> This is the upstream protocol doc, kept for provenance. Two things in it are wrong for
> our lineage and will silently corrupt results if followed literally:
>
> 1. **`action_horizon=10` (line ~59) is wrong.** The `pi05_ctc502_60k` checkpoint and all
>    three fine-tuned arms were trained at **50**. The "10" describes the transfer bundle's
>    unrelated base config.
> 2. **Norm stats.** This doc does not say which to use. The base must be served **z-score**
>    against its own flat ctc502 stats; the VACE / MIMICGEN / POSE6DAUG arms must be served
>    **quantile** against `ctc502_qnorm`. Serving either under `robocasa365_human300` is the
>    single biggest error we found.
>
> Measured impact of these two together: **1/8 -> 5/8** on identical seeded episodes,
> **7/160 -> 17/160** on the replay set.
>
> Also note: the **135/200 = 67.5% reference below is retracted by its own authors** and the
> 160-episode bank's 7/160 is not comparable to our numbers (different serving config,
> different noise scheme, different GPU architecture). See `EVAL_SETS.md`.
>
> What this doc IS still authoritative for: the per-episode seeding scheme
> (`env.reset(seed=ep_id)`, globals reseeded to `ep_id`, client-side noise
> `default_rng(action_seed + ep_id*1_000_003 + query_idx)`), the deterministic XLA flags,
> `replan_steps=5`, `resize_size=224`, camera 256, `action_dim=32`, and the fact that
> `--args.start_episode_idx` is a no-op superseded by `--args.env_seed_offset`.

---

# Evaluation seeds — pi0.5 ctc502 (RoboCasa PnP Counter→Cabinet)

**Revised 2026-09-22.** The previous version of this file described a seeding scheme
("PCG64 reset to 42 at the start of each episode, one [50,32] noise tensor per query") that the
shipped code did **not** implement. This revision documents what the code in `transfer.tar`
actually does after the reproducibility fix, and what was validated.

## What was wrong in the original bundle
1. **Environment history dependence.** `examples/robocasa/main.py` built the env once
   (`gym.make(..., seed=7)`) and called `env.reset()` unseeded. The scene of episode *i* therefore
   depended on how many resets happened before it in the same process (`--start_episode_idx` warm-up
   resets existed only to work around this). Two shards, or one interrupted run, do not see the
   same episodes.
2. **Policy noise history dependence, across clients.** The websocket policy server holds a single
   JAX PRNG (`Policy._rng`, split once per `infer()`). The flow-matching noise for a query depends on
   how many queries the server has answered so far — including queries from *other* clients
   sharing the server. Nothing in the client seeded it.
3. **Cross-process float nondeterminism.** With default XLA flags, two policy servers given the
   *same* observation and the *same* noise returned actions differing by up to 2.5e-2 (A100, same
   node). Cause: per-process XLA autotuning selecting different kernels.
4. The documented `action_horizon=50` does not match the shipped config
   (`pi05_robocasa_target_PickPlaceCounterToCabinet`: `action_horizon=10`), and the claimed
   per-episode PCG64 seeding did not exist anywhere in the code.

## What the fixed code does
Every episode has a global id `ep_id = env_seed_offset + episode_idx`, and every source of
randomness is keyed on it:
- **Env seed = `ep_id`**, applied with `env.reset(seed=ep_id)` on every episode
  (`gym_wrapper.reset(seed)` reseeds `env.rng`, which drives layout/style choice, object placement,
  robot init-pose noise and agentview camera noise). Episode *i* is identical regardless of which
  episodes ran before it, in which process, on which GPU.
- **Global NumPy / Python RNGs reseeded to `ep_id`** right before the reset. robosuite/robocasa
  still use `np.random` in a few places (observable corrupters, wrist-camera randomisation if
  enabled, some fixture colours); with the default task config none of them affect this task's
  state, but reseeding closes that path for other configs.
- **Action noise = client-generated**, `np.random.default_rng(action_seed + ep_id*1_000_003 + query_idx)`
  drawn as `[action_horizon, action_dim]` (read from the server's metadata: 10 × 32 for this
  config). Passed in the inference request as `"noise"`; `Policy.infer` validates the shape and
  feeds it to `sample_actions(noise=...)`. The server's internal RNG is never consulted when
  `noise` is supplied. The formula is the same one used in the GR00T exact-replay eval so the two
  pipelines share a convention.
- **Deterministic XLA** — `scripts/serve_policy.py` sets
  `XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0` (setdefault; override by
  exporting your own `XLA_FLAGS`).
- **Per-episode action traces** — `actions_{i}_{success|failure}.npy` (raw policy outputs before
  `convert_action`) next to the rollout mp4s, so two runs can be compared bit-for-bit.
- `--args.env_names PickPlaceCounterToCabinet` evaluates exactly one env instead of expanding a task set.
- `--args.start_episode_idx` is now a no-op (kept for CLI compatibility); use `--args.env_seed_offset`
  to select an episode range for sharding. Shards are now exactly reproducible: shard *k* covering
  episodes `[a, b)` runs `--args.env_seed_offset a --args.num_trials (b-a)`.

Known residual: MuJoCo EGL rendering is not bit-exact across processes — measured 1 pixel / 1 LSB in
the wrist image between two cold processes on the same GPU (agentview images identical, physics
state identical). Fed to the policy with identical noise this produced identical actions
(max|Δ|=0), and the PyAV mp4 round-trip is process-independent, so it did not affect the
validation below; it is noted because it is the one component outside the seeding scheme.

## Fixed eval settings
`replan_steps=5`, `resize_size=224`, camera 256, `action_horizon=10`, `action_dim=32`,
`mp4_roundtrip=True`, `generative_textures=False`, checkpoint kind = raw non-EMA params,
norm stats = `assets/pi0_fast_robocasa_pretrain_human300/robocasa365_human300/norm_stats.json`
(the quantile stats the config trains with — see INSTALL.md §4, the flat
`<ckpt>/assets/norm_stats.json` is *not* usable).

## Running
```bash
cd code/openpi
# server (one per GPU); XLA determinism flags are applied by the script
CUDA_VISIBLE_DEVICES=0 python scripts/serve_policy.py --port 8000 policy:checkpoint \
  --policy.config pi05_robocasa_target_PickPlaceCounterToCabinet --policy.dir <path>/pi05_ctc502_60k_rawparams
# client: episodes 0..199, env seeds 0..199, action seed 42
MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=0 CUDA_VISIBLE_DEVICES=0 python examples/robocasa/main.py \
  --args.host 127.0.0.1 --args.port 8000 --args.split target \
  --args.env_names PickPlaceCounterToCabinet --args.num_trials 200 \
  --args.env_seed_offset 0 --args.action_seed 42 --args.log_dir <out>
```
`MUJOCO_EGL_DEVICE_ID` must equal the CUDA device index you mask to.

## Validation (2026-09-22, A100 node, GPUs 2/3)
- **Env only** (`val/val_env.py`): episodes 5–9 started cold on one GPU vs. after episodes 0–4 on
  another: layout, style, object pose, base pose identical for all 5 → history-independent. PASS.
- **Policy** (`val/val_noise.py`): same obs + same client noise → bit-identical actions on repeat
  (max|Δ|=0) and across two servers on different GPUs (max|Δ|=0 with the XLA flags; 2.5e-2
  without). Without client noise, repeat queries differ (server RNG advances) — expected.
  Wrong-shape noise is rejected. PASS.
- **End-to-end** (`val/val_e2e.sh`): episodes 0–7 in one process vs. episodes 5–7 started cold in
  another process on another GPU: see `VALIDATION_RESULT` below.

VALIDATION_RESULT (main_optimized.py, 60k checkpoint, target split, action seed 42):
  ep5: SEQ failure | COLD failure | actions bit-identical for the first 25 steps (5 queries), then diverge
  ep6: SEQ failure | COLD failure | actions bit-identical for the first 30 steps (6 queries), then diverge
  ep7: SEQ failure | COLD failure | actions bit-identical for all 750 steps (150 queries), rollout mp4 identical
  Diagnosis of the ep5/ep6 divergence (val/val_replay.py): replaying the identical action sequence in
  two fresh processes gives identical physics state (qpos) at every step, but the rendered
  `agentview_right` / `eye_in_hand` images differ at a few pixels on some steps (GPU OpenGL/EGL
  rasterisation is not bit-reproducible across processes on this node; disabling shadows/reflection
  does not remove it). Once such a frame is fed to the policy the closed loop diverges. Everything
  the eval code controls (scene, robot init, placement, noise, model numerics) is exact; the
  remaining variation is the renderer. Runs are therefore reproducible up to this per-frame render
  jitter; report success rates over the full 200-episode set rather than per-episode outcomes.
  Note: 1/8 successes in this 8-episode check (eps 0-7) — far below the 67.5% previously claimed
  for this checkpoint; the previous number came from the unseeded code path and is not reproducible.

## Reference result
The previously reported **135 / 200 = 67.5%** for `pi05_ctc502_60k_rawparams` was produced with the
*original* history-dependent code and is **not** reproducible from the seeds it listed. Re-run with
the fixed code to obtain a reproducible number; per-episode `.npy` traces make any re-run checkable.

## Checkpoints in this repo
- `pi05_ctc502_60k_rawparams` — 60k non-EMA raw weights (ctc502_59999_rawparams)
- `pi05_ctc502_30k_rawparams` — 30k non-EMA raw weights (ctc502_repro29999_rawparams)
Both now carry `assets/robocasa365_human300/norm_stats.json` (copy of the config asset) so that
`serve_policy.py` loads them without edits.