eval_stack: document precedence — flag that upstream EVAL_SEEDS.md action_horizon=10 is wrong (should be 50) and state norm-stat rules
b963255 verified |
Download eval_stack/docs/EVAL_SEEDS.md from Ronaldo-GOAT/transfer: direct link, hf CLI and curl.
- Browser
- Download file 9.77 kB
-
https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/docs/EVAL_SEEDS.md
- Command line
-
hf download hf://Ronaldo-GOAT/transfer/eval_stack/docs/EVAL_SEEDS.md
-
curl -L -o EVAL_SEEDS.md https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/docs/EVAL_SEEDS.md
9.77 kB
| > **SUPERSEDED IN PART — read `../README.md` first.** | |
| > | |
| > This is the upstream protocol doc, kept for provenance. Two things in it are wrong for | |
| > our lineage and will silently corrupt results if followed literally: | |
| > | |
| > 1. **`action_horizon=10` (line ~59) is wrong.** The `pi05_ctc502_60k` checkpoint and all | |
| > three fine-tuned arms were trained at **50**. The "10" describes the transfer bundle's | |
| > unrelated base config. | |
| > 2. **Norm stats.** This doc does not say which to use. The base must be served **z-score** | |
| > against its own flat ctc502 stats; the VACE / MIMICGEN / POSE6DAUG arms must be served | |
| > **quantile** against `ctc502_qnorm`. Serving either under `robocasa365_human300` is the | |
| > single biggest error we found. | |
| > | |
| > Measured impact of these two together: **1/8 -> 5/8** on identical seeded episodes, | |
| > **7/160 -> 17/160** on the replay set. | |
| > | |
| > Also note: the **135/200 = 67.5% reference below is retracted by its own authors** and the | |
| > 160-episode bank's 7/160 is not comparable to our numbers (different serving config, | |
| > different noise scheme, different GPU architecture). See `EVAL_SETS.md`. | |
| > | |
| > What this doc IS still authoritative for: the per-episode seeding scheme | |
| > (`env.reset(seed=ep_id)`, globals reseeded to `ep_id`, client-side noise | |
| > `default_rng(action_seed + ep_id*1_000_003 + query_idx)`), the deterministic XLA flags, | |
| > `replan_steps=5`, `resize_size=224`, camera 256, `action_dim=32`, and the fact that | |
| > `--args.start_episode_idx` is a no-op superseded by `--args.env_seed_offset`. | |
| --- | |
| # Evaluation seeds — pi0.5 ctc502 (RoboCasa PnP Counter→Cabinet) | |
| **Revised 2026-09-22.** The previous version of this file described a seeding scheme | |
| ("PCG64 reset to 42 at the start of each episode, one [50,32] noise tensor per query") that the | |
| shipped code did **not** implement. This revision documents what the code in `transfer.tar` | |
| actually does after the reproducibility fix, and what was validated. | |
| ## What was wrong in the original bundle | |
| 1. **Environment history dependence.** `examples/robocasa/main.py` built the env once | |
| (`gym.make(..., seed=7)`) and called `env.reset()` unseeded. The scene of episode *i* therefore | |
| depended on how many resets happened before it in the same process (`--start_episode_idx` warm-up | |
| resets existed only to work around this). Two shards, or one interrupted run, do not see the | |
| same episodes. | |
| 2. **Policy noise history dependence, across clients.** The websocket policy server holds a single | |
| JAX PRNG (`Policy._rng`, split once per `infer()`). The flow-matching noise for a query depends on | |
| how many queries the server has answered so far — including queries from *other* clients | |
| sharing the server. Nothing in the client seeded it. | |
| 3. **Cross-process float nondeterminism.** With default XLA flags, two policy servers given the | |
| *same* observation and the *same* noise returned actions differing by up to 2.5e-2 (A100, same | |
| node). Cause: per-process XLA autotuning selecting different kernels. | |
| 4. The documented `action_horizon=50` does not match the shipped config | |
| (`pi05_robocasa_target_PickPlaceCounterToCabinet`: `action_horizon=10`), and the claimed | |
| per-episode PCG64 seeding did not exist anywhere in the code. | |
| ## What the fixed code does | |
| Every episode has a global id `ep_id = env_seed_offset + episode_idx`, and every source of | |
| randomness is keyed on it: | |
| - **Env seed = `ep_id`**, applied with `env.reset(seed=ep_id)` on every episode | |
| (`gym_wrapper.reset(seed)` reseeds `env.rng`, which drives layout/style choice, object placement, | |
| robot init-pose noise and agentview camera noise). Episode *i* is identical regardless of which | |
| episodes ran before it, in which process, on which GPU. | |
| - **Global NumPy / Python RNGs reseeded to `ep_id`** right before the reset. robosuite/robocasa | |
| still use `np.random` in a few places (observable corrupters, wrist-camera randomisation if | |
| enabled, some fixture colours); with the default task config none of them affect this task's | |
| state, but reseeding closes that path for other configs. | |
| - **Action noise = client-generated**, `np.random.default_rng(action_seed + ep_id*1_000_003 + query_idx)` | |
| drawn as `[action_horizon, action_dim]` (read from the server's metadata: 10 × 32 for this | |
| config). Passed in the inference request as `"noise"`; `Policy.infer` validates the shape and | |
| feeds it to `sample_actions(noise=...)`. The server's internal RNG is never consulted when | |
| `noise` is supplied. The formula is the same one used in the GR00T exact-replay eval so the two | |
| pipelines share a convention. | |
| - **Deterministic XLA** — `scripts/serve_policy.py` sets | |
| `XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0` (setdefault; override by | |
| exporting your own `XLA_FLAGS`). | |
| - **Per-episode action traces** — `actions_{i}_{success|failure}.npy` (raw policy outputs before | |
| `convert_action`) next to the rollout mp4s, so two runs can be compared bit-for-bit. | |
| - `--args.env_names PickPlaceCounterToCabinet` evaluates exactly one env instead of expanding a task set. | |
| - `--args.start_episode_idx` is now a no-op (kept for CLI compatibility); use `--args.env_seed_offset` | |
| to select an episode range for sharding. Shards are now exactly reproducible: shard *k* covering | |
| episodes `[a, b)` runs `--args.env_seed_offset a --args.num_trials (b-a)`. | |
| Known residual: MuJoCo EGL rendering is not bit-exact across processes — measured 1 pixel / 1 LSB in | |
| the wrist image between two cold processes on the same GPU (agentview images identical, physics | |
| state identical). Fed to the policy with identical noise this produced identical actions | |
| (max|Δ|=0), and the PyAV mp4 round-trip is process-independent, so it did not affect the | |
| validation below; it is noted because it is the one component outside the seeding scheme. | |
| ## Fixed eval settings | |
| `replan_steps=5`, `resize_size=224`, camera 256, `action_horizon=10`, `action_dim=32`, | |
| `mp4_roundtrip=True`, `generative_textures=False`, checkpoint kind = raw non-EMA params, | |
| norm stats = `assets/pi0_fast_robocasa_pretrain_human300/robocasa365_human300/norm_stats.json` | |
| (the quantile stats the config trains with — see INSTALL.md §4, the flat | |
| `<ckpt>/assets/norm_stats.json` is *not* usable). | |
| ## Running | |
| ```bash | |
| cd code/openpi | |
| # server (one per GPU); XLA determinism flags are applied by the script | |
| CUDA_VISIBLE_DEVICES=0 python scripts/serve_policy.py --port 8000 policy:checkpoint \ | |
| --policy.config pi05_robocasa_target_PickPlaceCounterToCabinet --policy.dir <path>/pi05_ctc502_60k_rawparams | |
| # client: episodes 0..199, env seeds 0..199, action seed 42 | |
| MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=0 CUDA_VISIBLE_DEVICES=0 python examples/robocasa/main.py \ | |
| --args.host 127.0.0.1 --args.port 8000 --args.split target \ | |
| --args.env_names PickPlaceCounterToCabinet --args.num_trials 200 \ | |
| --args.env_seed_offset 0 --args.action_seed 42 --args.log_dir <out> | |
| ``` | |
| `MUJOCO_EGL_DEVICE_ID` must equal the CUDA device index you mask to. | |
| ## Validation (2026-09-22, A100 node, GPUs 2/3) | |
| - **Env only** (`val/val_env.py`): episodes 5–9 started cold on one GPU vs. after episodes 0–4 on | |
| another: layout, style, object pose, base pose identical for all 5 → history-independent. PASS. | |
| - **Policy** (`val/val_noise.py`): same obs + same client noise → bit-identical actions on repeat | |
| (max|Δ|=0) and across two servers on different GPUs (max|Δ|=0 with the XLA flags; 2.5e-2 | |
| without). Without client noise, repeat queries differ (server RNG advances) — expected. | |
| Wrong-shape noise is rejected. PASS. | |
| - **End-to-end** (`val/val_e2e.sh`): episodes 0–7 in one process vs. episodes 5–7 started cold in | |
| another process on another GPU: see `VALIDATION_RESULT` below. | |
| VALIDATION_RESULT (main_optimized.py, 60k checkpoint, target split, action seed 42): | |
| ep5: SEQ failure | COLD failure | actions bit-identical for the first 25 steps (5 queries), then diverge | |
| ep6: SEQ failure | COLD failure | actions bit-identical for the first 30 steps (6 queries), then diverge | |
| ep7: SEQ failure | COLD failure | actions bit-identical for all 750 steps (150 queries), rollout mp4 identical | |
| Diagnosis of the ep5/ep6 divergence (val/val_replay.py): replaying the identical action sequence in | |
| two fresh processes gives identical physics state (qpos) at every step, but the rendered | |
| `agentview_right` / `eye_in_hand` images differ at a few pixels on some steps (GPU OpenGL/EGL | |
| rasterisation is not bit-reproducible across processes on this node; disabling shadows/reflection | |
| does not remove it). Once such a frame is fed to the policy the closed loop diverges. Everything | |
| the eval code controls (scene, robot init, placement, noise, model numerics) is exact; the | |
| remaining variation is the renderer. Runs are therefore reproducible up to this per-frame render | |
| jitter; report success rates over the full 200-episode set rather than per-episode outcomes. | |
| Note: 1/8 successes in this 8-episode check (eps 0-7) — far below the 67.5% previously claimed | |
| for this checkpoint; the previous number came from the unseeded code path and is not reproducible. | |
| ## Reference result | |
| The previously reported **135 / 200 = 67.5%** for `pi05_ctc502_60k_rawparams` was produced with the | |
| *original* history-dependent code and is **not** reproducible from the seeds it listed. Re-run with | |
| the fixed code to obtain a reproducible number; per-episode `.npy` traces make any re-run checkable. | |
| ## Checkpoints in this repo | |
| - `pi05_ctc502_60k_rawparams` — 60k non-EMA raw weights (ctc502_59999_rawparams) | |
| - `pi05_ctc502_30k_rawparams` — 30k non-EMA raw weights (ctc502_repro29999_rawparams) | |
| Both now carry `assets/robocasa365_human300/norm_stats.json` (copy of the config asset) so that | |
| `serve_policy.py` loads them without edits. | |