transfer / eval_stack /docs /EVAL_SEEDS.md
Ronaldo-GOAT's picture
eval_stack: document precedence — flag that upstream EVAL_SEEDS.md action_horizon=10 is wrong (should be 50) and state norm-stat rules
b963255 verified
|
Raw History Blame Contribute Delete
9.77 kB

SUPERSEDED IN PART — read ../README.md first.

This is the upstream protocol doc, kept for provenance. Two things in it are wrong for our lineage and will silently corrupt results if followed literally:

  1. action_horizon=10 (line ~59) is wrong. The pi05_ctc502_60k checkpoint and all three fine-tuned arms were trained at 50. The "10" describes the transfer bundle's unrelated base config.
  2. Norm stats. This doc does not say which to use. The base must be served z-score against its own flat ctc502 stats; the VACE / MIMICGEN / POSE6DAUG arms must be served quantile against ctc502_qnorm. Serving either under robocasa365_human300 is the single biggest error we found.

Measured impact of these two together: 1/8 -> 5/8 on identical seeded episodes, 7/160 -> 17/160 on the replay set.

Also note: the 135/200 = 67.5% reference below is retracted by its own authors and the 160-episode bank's 7/160 is not comparable to our numbers (different serving config, different noise scheme, different GPU architecture). See EVAL_SETS.md.

What this doc IS still authoritative for: the per-episode seeding scheme (env.reset(seed=ep_id), globals reseeded to ep_id, client-side noise default_rng(action_seed + ep_id*1_000_003 + query_idx)), the deterministic XLA flags, replan_steps=5, resize_size=224, camera 256, action_dim=32, and the fact that --args.start_episode_idx is a no-op superseded by --args.env_seed_offset.


Evaluation seeds — pi0.5 ctc502 (RoboCasa PnP Counter→Cabinet)

Revised 2026-09-22. The previous version of this file described a seeding scheme ("PCG64 reset to 42 at the start of each episode, one [50,32] noise tensor per query") that the shipped code did not implement. This revision documents what the code in transfer.tar actually does after the reproducibility fix, and what was validated.

What was wrong in the original bundle

  1. Environment history dependence. examples/robocasa/main.py built the env once (gym.make(..., seed=7)) and called env.reset() unseeded. The scene of episode i therefore depended on how many resets happened before it in the same process (--start_episode_idx warm-up resets existed only to work around this). Two shards, or one interrupted run, do not see the same episodes.
  2. Policy noise history dependence, across clients. The websocket policy server holds a single JAX PRNG (Policy._rng, split once per infer()). The flow-matching noise for a query depends on how many queries the server has answered so far — including queries from other clients sharing the server. Nothing in the client seeded it.
  3. Cross-process float nondeterminism. With default XLA flags, two policy servers given the same observation and the same noise returned actions differing by up to 2.5e-2 (A100, same node). Cause: per-process XLA autotuning selecting different kernels.
  4. The documented action_horizon=50 does not match the shipped config (pi05_robocasa_target_PickPlaceCounterToCabinet: action_horizon=10), and the claimed per-episode PCG64 seeding did not exist anywhere in the code.

What the fixed code does

Every episode has a global id ep_id = env_seed_offset + episode_idx, and every source of randomness is keyed on it:

  • Env seed = ep_id, applied with env.reset(seed=ep_id) on every episode (gym_wrapper.reset(seed) reseeds env.rng, which drives layout/style choice, object placement, robot init-pose noise and agentview camera noise). Episode i is identical regardless of which episodes ran before it, in which process, on which GPU.
  • Global NumPy / Python RNGs reseeded to ep_id right before the reset. robosuite/robocasa still use np.random in a few places (observable corrupters, wrist-camera randomisation if enabled, some fixture colours); with the default task config none of them affect this task's state, but reseeding closes that path for other configs.
  • Action noise = client-generated, np.random.default_rng(action_seed + ep_id*1_000_003 + query_idx) drawn as [action_horizon, action_dim] (read from the server's metadata: 10 × 32 for this config). Passed in the inference request as "noise"; Policy.infer validates the shape and feeds it to sample_actions(noise=...). The server's internal RNG is never consulted when noise is supplied. The formula is the same one used in the GR00T exact-replay eval so the two pipelines share a convention.
  • Deterministic XLA — scripts/serve_policy.py sets XLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0 (setdefault; override by exporting your own XLA_FLAGS).
  • Per-episode action traces — actions_{i}_{success|failure}.npy (raw policy outputs before convert_action) next to the rollout mp4s, so two runs can be compared bit-for-bit.
  • --args.env_names PickPlaceCounterToCabinet evaluates exactly one env instead of expanding a task set.
  • --args.start_episode_idx is now a no-op (kept for CLI compatibility); use --args.env_seed_offset to select an episode range for sharding. Shards are now exactly reproducible: shard k covering episodes [a, b) runs --args.env_seed_offset a --args.num_trials (b-a).

Known residual: MuJoCo EGL rendering is not bit-exact across processes — measured 1 pixel / 1 LSB in the wrist image between two cold processes on the same GPU (agentview images identical, physics state identical). Fed to the policy with identical noise this produced identical actions (max|Δ|=0), and the PyAV mp4 round-trip is process-independent, so it did not affect the validation below; it is noted because it is the one component outside the seeding scheme.

Fixed eval settings

replan_steps=5, resize_size=224, camera 256, action_horizon=10, action_dim=32, mp4_roundtrip=True, generative_textures=False, checkpoint kind = raw non-EMA params, norm stats = assets/pi0_fast_robocasa_pretrain_human300/robocasa365_human300/norm_stats.json (the quantile stats the config trains with — see INSTALL.md §4, the flat <ckpt>/assets/norm_stats.json is not usable).

Running

cd code/openpi
# server (one per GPU); XLA determinism flags are applied by the script
CUDA_VISIBLE_DEVICES=0 python scripts/serve_policy.py --port 8000 policy:checkpoint \
  --policy.config pi05_robocasa_target_PickPlaceCounterToCabinet --policy.dir <path>/pi05_ctc502_60k_rawparams
# client: episodes 0..199, env seeds 0..199, action seed 42
MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=0 CUDA_VISIBLE_DEVICES=0 python examples/robocasa/main.py \
  --args.host 127.0.0.1 --args.port 8000 --args.split target \
  --args.env_names PickPlaceCounterToCabinet --args.num_trials 200 \
  --args.env_seed_offset 0 --args.action_seed 42 --args.log_dir <out>

MUJOCO_EGL_DEVICE_ID must equal the CUDA device index you mask to.

Validation (2026-09-22, A100 node, GPUs 2/3)

  • Env only (val/val_env.py): episodes 5–9 started cold on one GPU vs. after episodes 0–4 on another: layout, style, object pose, base pose identical for all 5 → history-independent. PASS.
  • Policy (val/val_noise.py): same obs + same client noise → bit-identical actions on repeat (max|Δ|=0) and across two servers on different GPUs (max|Δ|=0 with the XLA flags; 2.5e-2 without). Without client noise, repeat queries differ (server RNG advances) — expected. Wrong-shape noise is rejected. PASS.
  • End-to-end (val/val_e2e.sh): episodes 0–7 in one process vs. episodes 5–7 started cold in another process on another GPU: see VALIDATION_RESULT below.

VALIDATION_RESULT (main_optimized.py, 60k checkpoint, target split, action seed 42): ep5: SEQ failure | COLD failure | actions bit-identical for the first 25 steps (5 queries), then diverge ep6: SEQ failure | COLD failure | actions bit-identical for the first 30 steps (6 queries), then diverge ep7: SEQ failure | COLD failure | actions bit-identical for all 750 steps (150 queries), rollout mp4 identical Diagnosis of the ep5/ep6 divergence (val/val_replay.py): replaying the identical action sequence in two fresh processes gives identical physics state (qpos) at every step, but the rendered agentview_right / eye_in_hand images differ at a few pixels on some steps (GPU OpenGL/EGL rasterisation is not bit-reproducible across processes on this node; disabling shadows/reflection does not remove it). Once such a frame is fed to the policy the closed loop diverges. Everything the eval code controls (scene, robot init, placement, noise, model numerics) is exact; the remaining variation is the renderer. Runs are therefore reproducible up to this per-frame render jitter; report success rates over the full 200-episode set rather than per-episode outcomes. Note: 1/8 successes in this 8-episode check (eps 0-7) — far below the 67.5% previously claimed for this checkpoint; the previous number came from the unseeded code path and is not reproducible.

Reference result

The previously reported 135 / 200 = 67.5% for pi05_ctc502_60k_rawparams was produced with the original history-dependent code and is not reproducible from the seeds it listed. Re-run with the fixed code to obtain a reproducible number; per-episode .npy traces make any re-run checkable.

Checkpoints in this repo

  • pi05_ctc502_60k_rawparams — 60k non-EMA raw weights (ctc502_59999_rawparams)
  • pi05_ctc502_30k_rawparams — 30k non-EMA raw weights (ctc502_repro29999_rawparams) Both now carry assets/robocasa365_human300/norm_stats.json (copy of the config asset) so that serve_policy.py loads them without edits.