Consensus Selection (CS) on LIBERO-LONG โ checkpoints + dataset
| file | role |
|---|---|
policy_ep-0250_sr-0.596.ckpt |
base OAT8 policy โ used for baseline and CS (CS needs no retraining) |
tokenizer_ep-0950_mse-0.002.ckpt |
frozen OAT tokenizer + action normalizer โ REQUIRED, do not rename |
policy_awr16_e100.ckpt |
CS-D (AWR single-forward distillation) |
data/libero10_N500.zarr.zip |
training zarr โ NOT needed for evaluation (only to re-collect / re-train CS-D) |
Placement (the tokenizer path is baked into the policy config)
oat/my_models/policy_ep-0250_sr-0.596.ckpt
oat/my_models/tokenizer_ep-0950_mse-0.002.ckpt
oat/my_models/policy_awr16_e100.ckpt
# optional, training only:
unzip libero10_N500.zarr.zip -d oat/data/libero/
Setup: uv sync + git submodule update --init --recursive.
LIBERO envs come from the libero package (bddl + init_states); raw demo HDF5 are not needed.
Run (config n_test=500 = 50 rollouts x 10 tasks; -n 5 => 2500 rollouts)
cd oat
# baseline (OAT8, single sample)
MUJOCO_GL=egl uv run scripts/eval_policy_sim.py -c my_models/policy_ep-0250_sr-0.596.ckpt \
-o eval_out/base -n 5 --entropy_threshold 0 --use_k_tokens 8
# CS-N (N in {4,8,16,32})
MUJOCO_GL=egl uv run scripts/eval_policy_sim.py -c my_models/policy_ep-0250_sr-0.596.ckpt \
-o eval_out/cs8 -n 5 --entropy_threshold 0 --use_k_tokens 8 --bon_free 8 --bon_signal vote
# CS-D (single forward, no sampling)
MUJOCO_GL=egl uv run scripts/eval_policy_sim.py -c my_models/policy_awr16_e100.ckpt \
-o eval_out/csd -n 5 --entropy_threshold 0 --use_k_tokens 8
Per-task SR is in eval_out/<run>/eval_log.json (keys {task}/mean_success_rate_mean);
overall in mean_success_rate_mean.
Expected: base 0.581 | CS-8 0.690 | CS-16 0.712 | CS-32 0.717 | CS-D ~0.684