Download eval_stack/docs/EVAL_SEEDS.md from Ronaldo-GOAT/transfer: direct link, hf CLI and curl.
- Browser
- Download file 9.77 kB
-
https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/docs/EVAL_SEEDS.md
- Command line
-
hf download hf://Ronaldo-GOAT/transfer/eval_stack/docs/EVAL_SEEDS.md
-
curl -L -o EVAL_SEEDS.md https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/eval_stack/docs/EVAL_SEEDS.md
SUPERSEDED IN PART — read
../README.mdfirst.This is the upstream protocol doc, kept for provenance. Two things in it are wrong for our lineage and will silently corrupt results if followed literally:
action_horizon=10(line ~59) is wrong. Thepi05_ctc502_60kcheckpoint and all three fine-tuned arms were trained at 50. The "10" describes the transfer bundle's unrelated base config.- Norm stats. This doc does not say which to use. The base must be served z-score against its own flat ctc502 stats; the VACE / MIMICGEN / POSE6DAUG arms must be served quantile against
ctc502_qnorm. Serving either underrobocasa365_human300is the single biggest error we found.Measured impact of these two together: 1/8 -> 5/8 on identical seeded episodes, 7/160 -> 17/160 on the replay set.
Also note: the 135/200 = 67.5% reference below is retracted by its own authors and the 160-episode bank's 7/160 is not comparable to our numbers (different serving config, different noise scheme, different GPU architecture). See
EVAL_SETS.md.What this doc IS still authoritative for: the per-episode seeding scheme (
env.reset(seed=ep_id), globals reseeded toep_id, client-side noisedefault_rng(action_seed + ep_id*1_000_003 + query_idx)), the deterministic XLA flags,replan_steps=5,resize_size=224, camera 256,action_dim=32, and the fact that--args.start_episode_idxis a no-op superseded by--args.env_seed_offset.
Evaluation seeds — pi0.5 ctc502 (RoboCasa PnP Counter→Cabinet)
Revised 2026-09-22. The previous version of this file described a seeding scheme
("PCG64 reset to 42 at the start of each episode, one [50,32] noise tensor per query") that the
shipped code did not implement. This revision documents what the code in transfer.tar
actually does after the reproducibility fix, and what was validated.
What was wrong in the original bundle
- Environment history dependence.
examples/robocasa/main.pybuilt the env once (gym.make(..., seed=7)) and calledenv.reset()unseeded. The scene of episode i therefore depended on how many resets happened before it in the same process (--start_episode_idxwarm-up resets existed only to work around this). Two shards, or one interrupted run, do not see the same episodes. - Policy noise history dependence, across clients. The websocket policy server holds a single
JAX PRNG (
Policy._rng, split once perinfer()). The flow-matching noise for a query depends on how many queries the server has answered so far — including queries from other clients sharing the server. Nothing in the client seeded it. - Cross-process float nondeterminism. With default XLA flags, two policy servers given the same observation and the same noise returned actions differing by up to 2.5e-2 (A100, same node). Cause: per-process XLA autotuning selecting different kernels.
- The documented
action_horizon=50does not match the shipped config (pi05_robocasa_target_PickPlaceCounterToCabinet:action_horizon=10), and the claimed per-episode PCG64 seeding did not exist anywhere in the code.
What the fixed code does
Every episode has a global id ep_id = env_seed_offset + episode_idx, and every source of
randomness is keyed on it:
- Env seed =
ep_id, applied withenv.reset(seed=ep_id)on every episode (gym_wrapper.reset(seed)reseedsenv.rng, which drives layout/style choice, object placement, robot init-pose noise and agentview camera noise). Episode i is identical regardless of which episodes ran before it, in which process, on which GPU. - Global NumPy / Python RNGs reseeded to
ep_idright before the reset. robosuite/robocasa still usenp.randomin a few places (observable corrupters, wrist-camera randomisation if enabled, some fixture colours); with the default task config none of them affect this task's state, but reseeding closes that path for other configs. - Action noise = client-generated,
np.random.default_rng(action_seed + ep_id*1_000_003 + query_idx)drawn as[action_horizon, action_dim](read from the server's metadata: 10 × 32 for this config). Passed in the inference request as"noise";Policy.infervalidates the shape and feeds it tosample_actions(noise=...). The server's internal RNG is never consulted whennoiseis supplied. The formula is the same one used in the GR00T exact-replay eval so the two pipelines share a convention. - Deterministic XLA —
scripts/serve_policy.pysetsXLA_FLAGS=--xla_gpu_deterministic_ops=true --xla_gpu_autotune_level=0(setdefault; override by exporting your ownXLA_FLAGS). - Per-episode action traces —
actions_{i}_{success|failure}.npy(raw policy outputs beforeconvert_action) next to the rollout mp4s, so two runs can be compared bit-for-bit. --args.env_names PickPlaceCounterToCabinetevaluates exactly one env instead of expanding a task set.--args.start_episode_idxis now a no-op (kept for CLI compatibility); use--args.env_seed_offsetto select an episode range for sharding. Shards are now exactly reproducible: shard k covering episodes[a, b)runs--args.env_seed_offset a --args.num_trials (b-a).
Known residual: MuJoCo EGL rendering is not bit-exact across processes — measured 1 pixel / 1 LSB in the wrist image between two cold processes on the same GPU (agentview images identical, physics state identical). Fed to the policy with identical noise this produced identical actions (max|Δ|=0), and the PyAV mp4 round-trip is process-independent, so it did not affect the validation below; it is noted because it is the one component outside the seeding scheme.
Fixed eval settings
replan_steps=5, resize_size=224, camera 256, action_horizon=10, action_dim=32,
mp4_roundtrip=True, generative_textures=False, checkpoint kind = raw non-EMA params,
norm stats = assets/pi0_fast_robocasa_pretrain_human300/robocasa365_human300/norm_stats.json
(the quantile stats the config trains with — see INSTALL.md §4, the flat
<ckpt>/assets/norm_stats.json is not usable).
Running
cd code/openpi
# server (one per GPU); XLA determinism flags are applied by the script
CUDA_VISIBLE_DEVICES=0 python scripts/serve_policy.py --port 8000 policy:checkpoint \
--policy.config pi05_robocasa_target_PickPlaceCounterToCabinet --policy.dir <path>/pi05_ctc502_60k_rawparams
# client: episodes 0..199, env seeds 0..199, action seed 42
MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=0 CUDA_VISIBLE_DEVICES=0 python examples/robocasa/main.py \
--args.host 127.0.0.1 --args.port 8000 --args.split target \
--args.env_names PickPlaceCounterToCabinet --args.num_trials 200 \
--args.env_seed_offset 0 --args.action_seed 42 --args.log_dir <out>
MUJOCO_EGL_DEVICE_ID must equal the CUDA device index you mask to.
Validation (2026-09-22, A100 node, GPUs 2/3)
- Env only (
val/val_env.py): episodes 5–9 started cold on one GPU vs. after episodes 0–4 on another: layout, style, object pose, base pose identical for all 5 → history-independent. PASS. - Policy (
val/val_noise.py): same obs + same client noise → bit-identical actions on repeat (max|Δ|=0) and across two servers on different GPUs (max|Δ|=0 with the XLA flags; 2.5e-2 without). Without client noise, repeat queries differ (server RNG advances) — expected. Wrong-shape noise is rejected. PASS. - End-to-end (
val/val_e2e.sh): episodes 0–7 in one process vs. episodes 5–7 started cold in another process on another GPU: seeVALIDATION_RESULTbelow.
VALIDATION_RESULT (main_optimized.py, 60k checkpoint, target split, action seed 42):
ep5: SEQ failure | COLD failure | actions bit-identical for the first 25 steps (5 queries), then diverge
ep6: SEQ failure | COLD failure | actions bit-identical for the first 30 steps (6 queries), then diverge
ep7: SEQ failure | COLD failure | actions bit-identical for all 750 steps (150 queries), rollout mp4 identical
Diagnosis of the ep5/ep6 divergence (val/val_replay.py): replaying the identical action sequence in
two fresh processes gives identical physics state (qpos) at every step, but the rendered
agentview_right / eye_in_hand images differ at a few pixels on some steps (GPU OpenGL/EGL
rasterisation is not bit-reproducible across processes on this node; disabling shadows/reflection
does not remove it). Once such a frame is fed to the policy the closed loop diverges. Everything
the eval code controls (scene, robot init, placement, noise, model numerics) is exact; the
remaining variation is the renderer. Runs are therefore reproducible up to this per-frame render
jitter; report success rates over the full 200-episode set rather than per-episode outcomes.
Note: 1/8 successes in this 8-episode check (eps 0-7) — far below the 67.5% previously claimed
for this checkpoint; the previous number came from the unseeded code path and is not reproducible.
Reference result
The previously reported 135 / 200 = 67.5% for pi05_ctc502_60k_rawparams was produced with the
original history-dependent code and is not reproducible from the seeds it listed. Re-run with
the fixed code to obtain a reproducible number; per-episode .npy traces make any re-run checkable.
Checkpoints in this repo
pi05_ctc502_60k_rawparams— 60k non-EMA raw weights (ctc502_59999_rawparams)pi05_ctc502_30k_rawparams— 30k non-EMA raw weights (ctc502_repro29999_rawparams) Both now carryassets/robocasa365_human300/norm_stats.json(copy of the config asset) so thatserve_policy.pyloads them without edits.