DiscoDemo-Stage1_RL-StackCube-alpha0

The data generator of DiscoDemo for the StackCube task (stack the red cube on the blue cube): a state-based reinforcement-learning policy trained in simulation — the P-RFCL baseline generator (α = 0: the same reverse-curriculum RL without the diversity reward). Rolling it out generates demonstrations such as those in the Stage2 dataset below.

Files

File Content
actor.pt Actor network weights (network_state_dict). Optimizer state is not included.
normalizer.pt Observation bounds used to normalize the policy input.
config.yaml Full training configuration (environment, curriculum, agent).

Only what is needed to roll the policy out is released; the critic and other training-only state are omitted.

Environment steps at this checkpoint: 50.0 M.

Usage

Loading requires the DiscoDemo training code: https://github.com/DAVIAN-Robotics/DiscoDemo.

Citation

@article{park2026discodemo,
  title   = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning},
  author  = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul},
  journal = {arXiv preprint},
  year    = {2026}
}
Downloads last month
4
Video Preview
loading

Collection including DAVIAN-Robotics/DiscoDemo-Stage1_RL-StackCube-alpha0