Image CCIL checkpoints
Selected PushT, Square and ToolHang policies
SR ± SE below uses pooled binary success over 150 tests: SE = sqrt(p(1-p)/150). Units are percentages. Each evaluation uses inference seeds 42, 1042, 2042 and 50 test episodes per seed; these are not independent training seeds. Repeated initial states may introduce correlation, which this binomial SE does not account for. PushT SR uses a 0.95 success threshold; continuous mean score is not SR.
| Task | Baseline | Sequential CCIL (BE) | E2E CCIL |
|---|---|---|---|
| PushT | 38.00 ± 3.96 | 51.33 ± 4.08, epoch780 | 58.00 ± 4.03, epoch500 |
| Square | 53.33 ± 4.07 | 61.33 ± 3.98, epoch380 | 70.00 ± 3.74, epoch420 |
| ToolHang | 23.33 ± 3.45 | 34.67 ± 3.89, epoch100 | 32.67 ± 3.83, epoch160 |
See machine-readable results and sources. Square E2E420 was evaluated with the corrected inference-seed handling on 2026-09-15. Other rows retain their published historical evaluations and were not rerun during this update; evaluation runtime/horizon differences remain possible. The old ToolHang baseline 36% number came from a single seed, not the standardized 3-by-50 summary.
Existing checkpoints, shared dynamics, and historical evaluation records are preserved. The table lists the selected main comparison only.
Other tasks
Can and Transport artifacts and published results are unchanged by this update.
Can (structured artifacts)
The Can files mirror the PushT and Square layouts. Standardized offline results
use diffusion inference seeds 42, 1042, and 2042; 50 test episodes per
inference seed; environment seeds 100000..100049; and success threshold
0.95. The baseline and sequential experiments use the first 20 demonstrations
from the 40-trajectory (20%) source dataset, corresponding to 10% of the full
dataset.
| Directory | Method | Checkpoint | Offline success rate |
|---|---|---|---|
can/baseline/ |
Diffusion-policy baseline used to initialize both sequential experiments | policy.pt (epoch 80) |
59/150 = 39.33% (46%, 38%, 34%) |
can/seq_be_epoch800/ |
Sequential CCIL, backward Euler | policy.pt (epoch 800) |
78/150 = 52.00% (56%, 48%, 52%) |
can/seq_noisy_action_noise0_epoch980/ |
Sequential CCIL, noisy action with zero action noise and 0.5 quantile filter | policy.pt (epoch 980) |
73/150 = 48.67% (58%, 48%, 40%) |
The two sequential policies share can/seq_shared_dynamics/dynamics.pkl and its
training configuration. Their distinct generation settings are stored as
augmentation_config.yaml in the respective policy directories. Each policy
directory contains exactly one selected checkpoint, its 56 fixed Robomimic
initial states, and the standardized 3-by-50 evaluation summary, JSON, and CSV.
The complete extraction, dynamics, augmentation, and finetuning recipe is kept
in can/run_can_seq.sh.
Transport (structured artifacts)
The Transport baseline and sequential experiment use the first 20 demonstrations from the 40-trajectory (20%) source dataset, corresponding to 10% of the full dataset. Standardized offline results use diffusion inference seeds 42, 1042, and 2042; 50 test episodes per inference seed; six train episodes retained in the evaluation batch; environment seeds 100000..100049; global scheduler RNG; render offsamples 0; and success threshold 0.95.
| Directory | Method | Checkpoint | Offline success rate |
|---|---|---|---|
transport/baseline/ |
Diffusion-policy baseline used to initialize sequential CCIL | policy.pt (epoch 440) |
76/150 = 50.67% (60%, 50%, 42%) |
transport/seq_be_epoch520/ |
Sequential CCIL, backward Euler | policy.pt (epoch 520) |
114/150 = 76.00% (84%, 72%, 72%) |
Each policy directory contains the selected checkpoint, its 56 fixed Robomimic initial states, and the standardized 3-by-50 evaluation files. The complete 20-demonstration sequential pipeline is stored in transport/run_transport_seq.sh.