> **SUPERSEDED IN PART — read `../README.md` first.** > > This is the upstream protocol doc, kept for provenance. Two things in it are wrong for > our lineage and will silently corrupt results if followed literally: > > 1. **`action_horizon=10` (line ~59) is wrong.** The `pi05_ctc502_60k` checkpoint and all > three fine-tuned arms were trained at **50**. The "10" describes the transfer bundle's > unrelated base config. > 2. **Norm stats.** This doc does not say which to use. The base must be served **z-score** > against its own flat ctc502 stats; the VACE / MIMICGEN / POSE6DAUG arms must be served > **quantile** against `ctc502_qnorm`. Serving either under `robocasa365_human300` is the > single biggest error we found. > > Measured impact of these two together: **1/8 -> 5/8** on identical seeded episodes, > **7/160 -> 17/160** on the replay set. > > Also note: the **135/200 = 67.5% reference below is retracted by its own authors** and the > 160-episode bank's 7/160 is not comparable to our numbers (different serving config, > different noise scheme, different GPU architecture). See `EVAL_SETS.md`. > > What this doc IS still authoritative for: the per-episode seeding scheme > (`env.reset(seed=ep_id)`, globals reseeded to `ep_id`, client-side noise > `default_rng(action_seed + ep_id*1_000_003 + query_idx)`), the deterministic XLA flags, > `replan_steps=5`, `resize_size=224`, camera 256, `action_dim=32`, and the fact that > `--args.start_episode_idx` is a no-op superseded by `--args.env_seed_offset`. --- # Evaluation seeds — pi0.5 ctc502 (RoboCasa PnP Counter→Cabinet) **Revised 2026-09-22.** The previous version of this file described a seeding scheme ("PCG64 reset to 42 at the start of each episode, one [50,32] noise tensor per query") that the shipped code did **not** implement. This revision documents what the code in `transfer.tar` actually does after the reproducibility fix, and what was validated. ## What was wrong in the original bundle 1. **Environment history dependence.** `examples/robocasa/main.py` built the env once (`gym.make(..., seed=7)`) and called `env.reset()` unseeded. The scene of episode *i* therefore depended on how many resets happened before it in the same process (`--start_episode_idx` warm-up resets existed only to work around this). Two shards, or one interrupted run, do not see the same episodes. 2. **Policy noise history dependence, across clients.** The websocket policy server holds a single JAX PRNG (`Policy._rng`, split once per `infer()`). The flow-matching noise for a query depends on how many queries the server has answered so far — including queries from *other* clients sharing the server. Nothing in the client seeded it. 3. **Cross-process float nondeterminism.** With default XLA flags, two policy servers given the *same* observation and the *same* noise returned actions differing by up to 2.5e-2 (A100, same node). Cause: per-process XLA autotuning selecting different kernels. 4. The documented `action_horizon=50` does not match the shipped config (`pi05_robocasa_target_PickPlaceCounterToCabinet`: `action_horizon=10`), and the claimed per-episode PCG64 seeding did not exist anywhere in the code. ## What the fixed code does Every episode has a global id `ep_id = env_seed_offset + episode_idx`, and every source of randomness is keyed on it: - **Env seed = `ep_id`**, applied with `env.reset(seed=ep_id)` on every episode (`gym_wrapper.reset(seed)` reseeds `env.rng`, which drives layout/style choice, object placement, robot init-pose noise and agentview camera noise). Episode *i* is identical regardless of which episodes ran before it, in which process, on which GPU. - **Global NumPy / Python RNGs reseeded to `ep_id`** right before the reset. robosuite/robocasa still use `np.random` in a few places (observable corrupters, wrist-camera randomisation if enabled, some fixture colours); with the default task config none of them affect this task's state, but reseeding closes that path for other configs. - **Action noise = client-generated**, `np.random.default_rng(action_seed + ep_id*1_000_003 + query_idx)` drawn as `[action_horizon, action_dim]` (read from the server's metadata: 10 × 32 for this config). Passed in the inference request as `"noise"`; `Policy.infer` validates the shape and feeds it to `sample_actions(noise=...)`. The server's internal RNG is never consulted when `noise` is supplied. The formula is the same one used in the GR00T exact-replay eval so the two pipelines share a convention. - **Deterministic XLA** — `scripts/serve_policy.py` sets `XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0` (setdefault; override by exporting your own `XLA_FLAGS`). - **Per-episode action traces** — `actions_{i}_{success|failure}.npy` (raw policy outputs before `convert_action`) next to the rollout mp4s, so two runs can be compared bit-for-bit. - `--args.env_names PickPlaceCounterToCabinet` evaluates exactly one env instead of expanding a task set. - `--args.start_episode_idx` is now a no-op (kept for CLI compatibility); use `--args.env_seed_offset` to select an episode range for sharding. Shards are now exactly reproducible: shard *k* covering episodes `[a, b)` runs `--args.env_seed_offset a --args.num_trials (b-a)`. Known residual: MuJoCo EGL rendering is not bit-exact across processes — measured 1 pixel / 1 LSB in the wrist image between two cold processes on the same GPU (agentview images identical, physics state identical). Fed to the policy with identical noise this produced identical actions (max|Δ|=0), and the PyAV mp4 round-trip is process-independent, so it did not affect the validation below; it is noted because it is the one component outside the seeding scheme. ## Fixed eval settings `replan_steps=5`, `resize_size=224`, camera 256, `action_horizon=10`, `action_dim=32`, `mp4_roundtrip=True`, `generative_textures=False`, checkpoint kind = raw non-EMA params, norm stats = `assets/pi0_fast_robocasa_pretrain_human300/robocasa365_human300/norm_stats.json` (the quantile stats the config trains with — see INSTALL.md §4, the flat `/assets/norm_stats.json` is *not* usable). ## Running ```bash cd code/openpi # server (one per GPU); XLA determinism flags are applied by the script CUDA_VISIBLE_DEVICES=0 python scripts/serve_policy.py --port 8000 policy:checkpoint \ --policy.config pi05_robocasa_target_PickPlaceCounterToCabinet --policy.dir /pi05_ctc502_60k_rawparams # client: episodes 0..199, env seeds 0..199, action seed 42 MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=0 CUDA_VISIBLE_DEVICES=0 python examples/robocasa/main.py \ --args.host 127.0.0.1 --args.port 8000 --args.split target \ --args.env_names PickPlaceCounterToCabinet --args.num_trials 200 \ --args.env_seed_offset 0 --args.action_seed 42 --args.log_dir ``` `MUJOCO_EGL_DEVICE_ID` must equal the CUDA device index you mask to. ## Validation (2026-09-22, A100 node, GPUs 2/3) - **Env only** (`val/val_env.py`): episodes 5–9 started cold on one GPU vs. after episodes 0–4 on another: layout, style, object pose, base pose identical for all 5 → history-independent. PASS. - **Policy** (`val/val_noise.py`): same obs + same client noise → bit-identical actions on repeat (max|Δ|=0) and across two servers on different GPUs (max|Δ|=0 with the XLA flags; 2.5e-2 without). Without client noise, repeat queries differ (server RNG advances) — expected. Wrong-shape noise is rejected. PASS. - **End-to-end** (`val/val_e2e.sh`): episodes 0–7 in one process vs. episodes 5–7 started cold in another process on another GPU: see `VALIDATION_RESULT` below. VALIDATION_RESULT (main_optimized.py, 60k checkpoint, target split, action seed 42): ep5: SEQ failure | COLD failure | actions bit-identical for the first 25 steps (5 queries), then diverge ep6: SEQ failure | COLD failure | actions bit-identical for the first 30 steps (6 queries), then diverge ep7: SEQ failure | COLD failure | actions bit-identical for all 750 steps (150 queries), rollout mp4 identical Diagnosis of the ep5/ep6 divergence (val/val_replay.py): replaying the identical action sequence in two fresh processes gives identical physics state (qpos) at every step, but the rendered `agentview_right` / `eye_in_hand` images differ at a few pixels on some steps (GPU OpenGL/EGL rasterisation is not bit-reproducible across processes on this node; disabling shadows/reflection does not remove it). Once such a frame is fed to the policy the closed loop diverges. Everything the eval code controls (scene, robot init, placement, noise, model numerics) is exact; the remaining variation is the renderer. Runs are therefore reproducible up to this per-frame render jitter; report success rates over the full 200-episode set rather than per-episode outcomes. Note: 1/8 successes in this 8-episode check (eps 0-7) — far below the 67.5% previously claimed for this checkpoint; the previous number came from the unseeded code path and is not reproducible. ## Reference result The previously reported **135 / 200 = 67.5%** for `pi05_ctc502_60k_rawparams` was produced with the *original* history-dependent code and is **not** reproducible from the seeds it listed. Re-run with the fixed code to obtain a reproducible number; per-episode `.npy` traces make any re-run checkable. ## Checkpoints in this repo - `pi05_ctc502_60k_rawparams` — 60k non-EMA raw weights (ctc502_59999_rawparams) - `pi05_ctc502_30k_rawparams` — 30k non-EMA raw weights (ctc502_repro29999_rawparams) Both now carry `assets/robocasa365_human300/norm_stats.json` (copy of the config asset) so that `serve_policy.py` loads them without edits.