transfer / eval_stack /docs /EVAL_SEEDS.md
Ronaldo-GOAT's picture
eval_stack: document precedence — flag that upstream EVAL_SEEDS.md action_horizon=10 is wrong (should be 50) and state norm-stat rules
b963255 verified
|
Raw History Blame Contribute Delete
9.77 kB
> **SUPERSEDED IN PART — read `../README.md` first.**
>
> This is the upstream protocol doc, kept for provenance. Two things in it are wrong for
> our lineage and will silently corrupt results if followed literally:
>
> 1. **`action_horizon=10` (line ~59) is wrong.** The `pi05_ctc502_60k` checkpoint and all
> three fine-tuned arms were trained at **50**. The "10" describes the transfer bundle's
> unrelated base config.
> 2. **Norm stats.** This doc does not say which to use. The base must be served **z-score**
> against its own flat ctc502 stats; the VACE / MIMICGEN / POSE6DAUG arms must be served
> **quantile** against `ctc502_qnorm`. Serving either under `robocasa365_human300` is the
> single biggest error we found.
>
> Measured impact of these two together: **1/8 -> 5/8** on identical seeded episodes,
> **7/160 -> 17/160** on the replay set.
>
> Also note: the **135/200 = 67.5% reference below is retracted by its own authors** and the
> 160-episode bank's 7/160 is not comparable to our numbers (different serving config,
> different noise scheme, different GPU architecture). See `EVAL_SETS.md`.
>
> What this doc IS still authoritative for: the per-episode seeding scheme
> (`env.reset(seed=ep_id)`, globals reseeded to `ep_id`, client-side noise
> `default_rng(action_seed + ep_id*1_000_003 + query_idx)`), the deterministic XLA flags,
> `replan_steps=5`, `resize_size=224`, camera 256, `action_dim=32`, and the fact that
> `--args.start_episode_idx` is a no-op superseded by `--args.env_seed_offset`.
---
# Evaluation seeds — pi0.5 ctc502 (RoboCasa PnP Counter→Cabinet)
**Revised 2026-09-22.** The previous version of this file described a seeding scheme
("PCG64 reset to 42 at the start of each episode, one [50,32] noise tensor per query") that the
shipped code did **not** implement. This revision documents what the code in `transfer.tar`
actually does after the reproducibility fix, and what was validated.
## What was wrong in the original bundle
1. **Environment history dependence.** `examples/robocasa/main.py` built the env once
(`gym.make(..., seed=7)`) and called `env.reset()` unseeded. The scene of episode *i* therefore
depended on how many resets happened before it in the same process (`--start_episode_idx` warm-up
resets existed only to work around this). Two shards, or one interrupted run, do not see the
same episodes.
2. **Policy noise history dependence, across clients.** The websocket policy server holds a single
JAX PRNG (`Policy._rng`, split once per `infer()`). The flow-matching noise for a query depends on
how many queries the server has answered so far — including queries from *other* clients
sharing the server. Nothing in the client seeded it.
3. **Cross-process float nondeterminism.** With default XLA flags, two policy servers given the
*same* observation and the *same* noise returned actions differing by up to 2.5e-2 (A100, same
node). Cause: per-process XLA autotuning selecting different kernels.
4. The documented `action_horizon=50` does not match the shipped config
(`pi05_robocasa_target_PickPlaceCounterToCabinet`: `action_horizon=10`), and the claimed
per-episode PCG64 seeding did not exist anywhere in the code.
## What the fixed code does
Every episode has a global id `ep_id = env_seed_offset + episode_idx`, and every source of
randomness is keyed on it:
- **Env seed = `ep_id`**, applied with `env.reset(seed=ep_id)` on every episode
(`gym_wrapper.reset(seed)` reseeds `env.rng`, which drives layout/style choice, object placement,
robot init-pose noise and agentview camera noise). Episode *i* is identical regardless of which
episodes ran before it, in which process, on which GPU.
- **Global NumPy / Python RNGs reseeded to `ep_id`** right before the reset. robosuite/robocasa
still use `np.random` in a few places (observable corrupters, wrist-camera randomisation if
enabled, some fixture colours); with the default task config none of them affect this task's
state, but reseeding closes that path for other configs.
- **Action noise = client-generated**, `np.random.default_rng(action_seed + ep_id*1_000_003 + query_idx)`
drawn as `[action_horizon, action_dim]` (read from the server's metadata: 10 × 32 for this
config). Passed in the inference request as `"noise"`; `Policy.infer` validates the shape and
feeds it to `sample_actions(noise=...)`. The server's internal RNG is never consulted when
`noise` is supplied. The formula is the same one used in the GR00T exact-replay eval so the two
pipelines share a convention.
- **Deterministic XLA** — `scripts/serve_policy.py` sets
`XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0` (setdefault; override by
exporting your own `XLA_FLAGS`).
- **Per-episode action traces** — `actions_{i}_{success|failure}.npy` (raw policy outputs before
`convert_action`) next to the rollout mp4s, so two runs can be compared bit-for-bit.
- `--args.env_names PickPlaceCounterToCabinet` evaluates exactly one env instead of expanding a task set.
- `--args.start_episode_idx` is now a no-op (kept for CLI compatibility); use `--args.env_seed_offset`
to select an episode range for sharding. Shards are now exactly reproducible: shard *k* covering
episodes `[a, b)` runs `--args.env_seed_offset a --args.num_trials (b-a)`.
Known residual: MuJoCo EGL rendering is not bit-exact across processes — measured 1 pixel / 1 LSB in
the wrist image between two cold processes on the same GPU (agentview images identical, physics
state identical). Fed to the policy with identical noise this produced identical actions
(max|Δ|=0), and the PyAV mp4 round-trip is process-independent, so it did not affect the
validation below; it is noted because it is the one component outside the seeding scheme.
## Fixed eval settings
`replan_steps=5`, `resize_size=224`, camera 256, `action_horizon=10`, `action_dim=32`,
`mp4_roundtrip=True`, `generative_textures=False`, checkpoint kind = raw non-EMA params,
norm stats = `assets/pi0_fast_robocasa_pretrain_human300/robocasa365_human300/norm_stats.json`
(the quantile stats the config trains with — see INSTALL.md §4, the flat
`<ckpt>/assets/norm_stats.json` is *not* usable).
## Running
```bash
cd code/openpi
# server (one per GPU); XLA determinism flags are applied by the script
CUDA_VISIBLE_DEVICES=0 python scripts/serve_policy.py --port 8000 policy:checkpoint \
--policy.config pi05_robocasa_target_PickPlaceCounterToCabinet --policy.dir <path>/pi05_ctc502_60k_rawparams
# client: episodes 0..199, env seeds 0..199, action seed 42
MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=0 CUDA_VISIBLE_DEVICES=0 python examples/robocasa/main.py \
--args.host 127.0.0.1 --args.port 8000 --args.split target \
--args.env_names PickPlaceCounterToCabinet --args.num_trials 200 \
--args.env_seed_offset 0 --args.action_seed 42 --args.log_dir <out>
```
`MUJOCO_EGL_DEVICE_ID` must equal the CUDA device index you mask to.
## Validation (2026-09-22, A100 node, GPUs 2/3)
- **Env only** (`val/val_env.py`): episodes 5–9 started cold on one GPU vs. after episodes 0–4 on
another: layout, style, object pose, base pose identical for all 5 → history-independent. PASS.
- **Policy** (`val/val_noise.py`): same obs + same client noise → bit-identical actions on repeat
(max|Δ|=0) and across two servers on different GPUs (max|Δ|=0 with the XLA flags; 2.5e-2
without). Without client noise, repeat queries differ (server RNG advances) — expected.
Wrong-shape noise is rejected. PASS.
- **End-to-end** (`val/val_e2e.sh`): episodes 0–7 in one process vs. episodes 5–7 started cold in
another process on another GPU: see `VALIDATION_RESULT` below.
VALIDATION_RESULT (main_optimized.py, 60k checkpoint, target split, action seed 42):
ep5: SEQ failure | COLD failure | actions bit-identical for the first 25 steps (5 queries), then diverge
ep6: SEQ failure | COLD failure | actions bit-identical for the first 30 steps (6 queries), then diverge
ep7: SEQ failure | COLD failure | actions bit-identical for all 750 steps (150 queries), rollout mp4 identical
Diagnosis of the ep5/ep6 divergence (val/val_replay.py): replaying the identical action sequence in
two fresh processes gives identical physics state (qpos) at every step, but the rendered
`agentview_right` / `eye_in_hand` images differ at a few pixels on some steps (GPU OpenGL/EGL
rasterisation is not bit-reproducible across processes on this node; disabling shadows/reflection
does not remove it). Once such a frame is fed to the policy the closed loop diverges. Everything
the eval code controls (scene, robot init, placement, noise, model numerics) is exact; the
remaining variation is the renderer. Runs are therefore reproducible up to this per-frame render
jitter; report success rates over the full 200-episode set rather than per-episode outcomes.
Note: 1/8 successes in this 8-episode check (eps 0-7) — far below the 67.5% previously claimed
for this checkpoint; the previous number came from the unseeded code path and is not reproducible.
## Reference result
The previously reported **135 / 200 = 67.5%** for `pi05_ctc502_60k_rawparams` was produced with the
*original* history-dependent code and is **not** reproducible from the seeds it listed. Re-run with
the fixed code to obtain a reproducible number; per-episode `.npy` traces make any re-run checkable.
## Checkpoints in this repo
- `pi05_ctc502_60k_rawparams` — 60k non-EMA raw weights (ctc502_59999_rawparams)
- `pi05_ctc502_30k_rawparams` — 30k non-EMA raw weights (ctc502_repro29999_rawparams)
Both now carry `assets/robocasa365_human300/norm_stats.json` (copy of the config asset) so that
`serve_policy.py` loads them without edits.