BOSS — pre-trained skill policies

Skill policies for the three Behavioral Cloning baselines of BOSS: Benchmark for Observation Space Shift in Long-Horizon Task (Yang et al., IEEE Robotics and Automation Letters, 2025), so the benchmark can be run without re-training.

Code: https://github.com/YY-GX/BOSS · Project page: https://boss-benchmark.github.io/

Contents

Folder Baseline Files
BCRNNPolicy_seed10000/run_001/ BC-RESNET-RNN task0..task43_model.pth
BCTransformerPolicy_seed10000/run_001/ BC-RESNET-T task0..task43_model.pth
BCViLTPolicy_seed10000/run_001/ BC-VIT-T task0..task43_model.pth

One checkpoint per skill, indexed by position in the boss_44 task list. All trained with seed=10000, observation space = third-person camera + 7-DoF joint angles + 2-DoF gripper state (no wrist camera), matching Section V-A of the paper.

Usage

huggingface-cli download yygx/BOSS-checkpoints --local-dir experiments/boss_44/0.0.0

python libero/lifelong/eval_skills_unaffected_by_oss.py \
  --benchmark boss_44 --seed 10000 --max_steps 400 \
  --model_path_folder experiments/boss_44/0.0.0/BCTransformerPolicy_seed10000/run_001/

Important: --max_steps 400

Evaluation hyper-parameters are read from the checkpoint, and these checkpoints carry max_steps=600 while the published results used 400. On BOSS-C3 task 1, BC-RESNET-T reaches a Delta to Upper Bound Ratio of 89% at 400 steps — the value in Figure 5 — and 67% at 600. Always pass --max_steps 400 to reproduce the paper.

Citation

@article{yang2025boss,
  title={BOSS: Benchmark for observation space shift in long-horizon task},
  author={Yang, Yue and Zhao, Linfeng and Ding, Mingyu and Bertasius, Gedas and Szafir, Daniel},
  journal={IEEE Robotics and Automation Letters},
  volume={10},
  number={9},
  pages={8882--8889},
  year={2025},
  publisher={IEEE}
}

Built on LIBERO (Liu et al., 2023).

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Paper for yygx/BOSS-checkpoints