Image CCIL checkpoints

Selected PushT, Square and ToolHang policies

SR ± SE below uses pooled binary success over 150 tests: SE = sqrt(p(1-p)/150). Units are percentages. Each evaluation uses inference seeds 42, 1042, 2042 and 50 test episodes per seed; these are not independent training seeds. Repeated initial states may introduce correlation, which this binomial SE does not account for. PushT SR uses a 0.95 success threshold; continuous mean score is not SR.

See machine-readable results and sources. Square E2E420 was evaluated with the corrected inference-seed handling on 2026-09-15. Other rows retain their published historical evaluations and were not rerun during this update; evaluation runtime/horizon differences remain possible. The old ToolHang baseline 36% number came from a single seed, not the standardized 3-by-50 summary.

Existing checkpoints, shared dynamics, and historical evaluation records are preserved. The table lists the selected main comparison only.

Other tasks

Can and Transport artifacts and published results are unchanged by this update.

Can (structured artifacts)

The Can files mirror the PushT and Square layouts. Standardized offline results use diffusion inference seeds 42, 1042, and 2042; 50 test episodes per inference seed; environment seeds 100000..100049; and success threshold 0.95. The baseline and sequential experiments use the first 20 demonstrations from the 40-trajectory (20%) source dataset, corresponding to 10% of the full dataset.

Directory Method Checkpoint Offline success rate
can/baseline/ Diffusion-policy baseline used to initialize both sequential experiments policy.pt (epoch 80) 59/150 = 39.33% (46%, 38%, 34%)
can/seq_be_epoch800/ Sequential CCIL, backward Euler policy.pt (epoch 800) 78/150 = 52.00% (56%, 48%, 52%)
can/seq_noisy_action_noise0_epoch980/ Sequential CCIL, noisy action with zero action noise and 0.5 quantile filter policy.pt (epoch 980) 73/150 = 48.67% (58%, 48%, 40%)

The two sequential policies share can/seq_shared_dynamics/dynamics.pkl and its training configuration. Their distinct generation settings are stored as augmentation_config.yaml in the respective policy directories. Each policy directory contains exactly one selected checkpoint, its 56 fixed Robomimic initial states, and the standardized 3-by-50 evaluation summary, JSON, and CSV. The complete extraction, dynamics, augmentation, and finetuning recipe is kept in can/run_can_seq.sh.

Transport (structured artifacts)

The Transport baseline and sequential experiment use the first 20 demonstrations from the 40-trajectory (20%) source dataset, corresponding to 10% of the full dataset. Standardized offline results use diffusion inference seeds 42, 1042, and 2042; 50 test episodes per inference seed; six train episodes retained in the evaluation batch; environment seeds 100000..100049; global scheduler RNG; render offsamples 0; and success threshold 0.95.

Directory Method Checkpoint Offline success rate
transport/baseline/ Diffusion-policy baseline used to initialize sequential CCIL policy.pt (epoch 440) 76/150 = 50.67% (60%, 50%, 42%)
transport/seq_be_epoch520/ Sequential CCIL, backward Euler policy.pt (epoch 520) 114/150 = 76.00% (84%, 72%, 72%)

Each policy directory contains the selected checkpoint, its 56 fixed Robomimic initial states, and the standardized 3-by-50 evaluation files. The complete 20-demonstration sequential pipeline is stored in transport/run_transport_seq.sh.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support