BOSS — pre-trained skill policies
Skill policies for the three Behavioral Cloning baselines of BOSS: Benchmark for Observation Space Shift in Long-Horizon Task (Yang et al., IEEE Robotics and Automation Letters, 2025), so the benchmark can be run without re-training.
Code: https://github.com/YY-GX/BOSS · Project page: https://boss-benchmark.github.io/
Contents
| Folder | Baseline | Files |
|---|---|---|
BCRNNPolicy_seed10000/run_001/ |
BC-RESNET-RNN | task0..task43_model.pth |
BCTransformerPolicy_seed10000/run_001/ |
BC-RESNET-T | task0..task43_model.pth |
BCViLTPolicy_seed10000/run_001/ |
BC-VIT-T | task0..task43_model.pth |
One checkpoint per skill, indexed by position in the boss_44 task list. All
trained with seed=10000, observation space = third-person camera + 7-DoF joint
angles + 2-DoF gripper state (no wrist camera), matching Section V-A of the paper.
Usage
huggingface-cli download yygx/BOSS-checkpoints --local-dir experiments/boss_44/0.0.0
python libero/lifelong/eval_skills_unaffected_by_oss.py \
--benchmark boss_44 --seed 10000 --max_steps 400 \
--model_path_folder experiments/boss_44/0.0.0/BCTransformerPolicy_seed10000/run_001/
Important: --max_steps 400
Evaluation hyper-parameters are read from the checkpoint, and these checkpoints
carry max_steps=600 while the published results used 400. On BOSS-C3 task 1,
BC-RESNET-T reaches a Delta to Upper Bound Ratio of 89% at 400 steps — the value
in Figure 5 — and 67% at 600. Always pass --max_steps 400 to reproduce the paper.
Citation
@article{yang2025boss,
title={BOSS: Benchmark for observation space shift in long-horizon task},
author={Yang, Yue and Zhao, Linfeng and Ding, Mingyu and Bertasius, Gedas and Szafir, Daniel},
journal={IEEE Robotics and Automation Letters},
volume={10},
number={9},
pages={8882--8889},
year={2025},
publisher={IEEE}
}
Built on LIBERO (Liu et al., 2023).