AdaptDICE trained models

Trained AdaptDICE target-domain models from Semi-Supervised Cross-Domain Imitation Learning. Paper: arXiv:2602.10793. Code: NYCU-RL-Bandits-Lab/CDIL. Datasets: rl-bandit-lab/CDIL.

The training inputs are already in the GitHub repo, so this repo does not duplicate them:

  • the source-domain DemoDICE models in pretrained_models/*.pickle
  • the pre-trained normalizing flows in flow_model/<env>/

Contents

adaptdice/<Env>/<set>_E<e>_I<a>-<b>/seed<k>.safetensors is the final checkpoint (iteration 500k) for seeds 0–4.

  • E<e>: number of labeled target expert trajectories (--expert_num_traj)
  • I<a>-<b>: unlabeled imperfect data, made of a expert + b random trajectories (--imperfect_dataset_default_info)
  • set1 = Default, set2 = Expert-Rich, set3 = Sub-Optimal-Rich
Env set1 set2 set3
Hopper E1_I10-50 E5_I20-50 E1_I50-100
HalfCheetah / Ant / Lift / Door E1_I1-100 E5_I1-100 E1_I5-500
Wipe E1_I10-50 E5_I10-50 E1_I50-100

Each file contains only network weights (no optimizer state). Keys are prefixed by the module name on agent.avatar_dice.Avatar:

  • actor.*: the target-domain policy
  • decoder.*: the state mapping G (before the flow)
  • action_decoder.*: the action mapping H (before the flow)
  • cost.* and critic.*

All networks use hidden size 256. Wipe was trained without the normalizing flow. All other envs used flow_model/<env>/ from the code repo.

Loading

Build the Avatar imitator the same way train_il.py does for --algorithm=avatar_dice, then load the weights:

from safetensors.torch import load_file

sd = load_file("adaptdice/Ant/set1_E1_I1-100/seed0.safetensors")
for m in ["actor", "cost", "critic", "decoder", "action_decoder"]:
    getattr(imitator, m).load_state_dict({k[len(m) + 1:]: v for k, v in sd.items() if k.startswith(m + ".")})

The policy expects normalized observations with a trailing 0 for the absorbing-state flag. train_il.py computes the normalization from the imperfect target-domain data (shift = -mean, scale = 1 / (std + 1e-3)), so use the same dataset files and counts listed above.

Citation

@misc{chu2026semi,
  title         = {Semi-Supervised Cross-Domain Imitation Learning},
  author        = {Chu, Li-Min and Ma, Kai-Siang and Chen, Ming-Hong and Hsieh, Ping-Chun},
  year          = {2026},
  eprint        = {2602.10793},
  archivePrefix = {arXiv}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for rl-bandit-lab/CDIL