AdaptDICE trained models
Trained AdaptDICE target-domain models from Semi-Supervised Cross-Domain Imitation Learning. Paper: arXiv:2602.10793. Code: NYCU-RL-Bandits-Lab/CDIL. Datasets: rl-bandit-lab/CDIL.
The training inputs are already in the GitHub repo, so this repo does not duplicate them:
- the source-domain DemoDICE models in
pretrained_models/*.pickle - the pre-trained normalizing flows in
flow_model/<env>/
Contents
adaptdice/<Env>/<set>_E<e>_I<a>-<b>/seed<k>.safetensors is the final checkpoint (iteration 500k) for seeds 0–4.
E<e>: number of labeled target expert trajectories (--expert_num_traj)I<a>-<b>: unlabeled imperfect data, made ofaexpert +brandom trajectories (--imperfect_dataset_default_info)set1= Default,set2= Expert-Rich,set3= Sub-Optimal-Rich
| Env | set1 | set2 | set3 |
|---|---|---|---|
| Hopper | E1_I10-50 | E5_I20-50 | E1_I50-100 |
| HalfCheetah / Ant / Lift / Door | E1_I1-100 | E5_I1-100 | E1_I5-500 |
| Wipe | E1_I10-50 | E5_I10-50 | E1_I50-100 |
Each file contains only network weights (no optimizer state). Keys are prefixed by the module name on agent.avatar_dice.Avatar:
actor.*: the target-domain policydecoder.*: the state mapping G (before the flow)action_decoder.*: the action mapping H (before the flow)cost.*andcritic.*
All networks use hidden size 256. Wipe was trained without the normalizing flow. All other envs used flow_model/<env>/ from the code repo.
Loading
Build the Avatar imitator the same way train_il.py does for --algorithm=avatar_dice, then load the weights:
from safetensors.torch import load_file
sd = load_file("adaptdice/Ant/set1_E1_I1-100/seed0.safetensors")
for m in ["actor", "cost", "critic", "decoder", "action_decoder"]:
getattr(imitator, m).load_state_dict({k[len(m) + 1:]: v for k, v in sd.items() if k.startswith(m + ".")})
The policy expects normalized observations with a trailing 0 for the absorbing-state flag. train_il.py computes the normalization from the imperfect target-domain data (shift = -mean, scale = 1 / (std + 1e-3)), so use the same dataset files and counts listed above.
Citation
@misc{chu2026semi,
title = {Semi-Supervised Cross-Domain Imitation Learning},
author = {Chu, Li-Min and Ma, Kai-Siang and Chen, Ming-Hong and Hsieh, Ping-Chun},
year = {2026},
eprint = {2602.10793},
archivePrefix = {arXiv}
}