Consensus Selection (CS) on LIBERO-LONG โ€” checkpoints + dataset

file role
policy_ep-0250_sr-0.596.ckpt base OAT8 policy โ€” used for baseline and CS (CS needs no retraining)
tokenizer_ep-0950_mse-0.002.ckpt frozen OAT tokenizer + action normalizer โ€” REQUIRED, do not rename
policy_awr16_e100.ckpt CS-D (AWR single-forward distillation)
data/libero10_N500.zarr.zip training zarr โ€” NOT needed for evaluation (only to re-collect / re-train CS-D)

Placement (the tokenizer path is baked into the policy config)

oat/my_models/policy_ep-0250_sr-0.596.ckpt
oat/my_models/tokenizer_ep-0950_mse-0.002.ckpt
oat/my_models/policy_awr16_e100.ckpt
# optional, training only:
unzip libero10_N500.zarr.zip -d oat/data/libero/

Setup: uv sync + git submodule update --init --recursive. LIBERO envs come from the libero package (bddl + init_states); raw demo HDF5 are not needed.

Run (config n_test=500 = 50 rollouts x 10 tasks; -n 5 => 2500 rollouts)

cd oat
# baseline (OAT8, single sample)
MUJOCO_GL=egl uv run scripts/eval_policy_sim.py -c my_models/policy_ep-0250_sr-0.596.ckpt \
  -o eval_out/base -n 5 --entropy_threshold 0 --use_k_tokens 8

# CS-N  (N in {4,8,16,32})
MUJOCO_GL=egl uv run scripts/eval_policy_sim.py -c my_models/policy_ep-0250_sr-0.596.ckpt \
  -o eval_out/cs8 -n 5 --entropy_threshold 0 --use_k_tokens 8 --bon_free 8 --bon_signal vote

# CS-D (single forward, no sampling)
MUJOCO_GL=egl uv run scripts/eval_policy_sim.py -c my_models/policy_awr16_e100.ckpt \
  -o eval_out/csd -n 5 --entropy_threshold 0 --use_k_tokens 8

Per-task SR is in eval_out/<run>/eval_log.json (keys {task}/mean_success_rate_mean); overall in mean_success_rate_mean.

Expected: base 0.581 | CS-8 0.690 | CS-16 0.712 | CS-32 0.717 | CS-D ~0.684

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading