DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

Anonymous release for ICLR 2027 submission

DistillAlign aligns and balances the mode-covering and mode-seeking objectives of the multi-stage video distillation pipeline — using only a 1.3B DMD teacher, it already surpasses baselines refined with a 14B DMD teacher.

Checkpoints

All released generators are Wan2.1-1.3B students; the size in each name refers to the teacher used during training.

Model Checkpoint Description
Initializer (1.3B teacher) distillalign_init_1p3b_teacher.pt Pre-DMD initializer, trained toward a Wan2.1-T2V-1.3B teacher
Initializer (14B teacher) distillalign_init_14b_teacher.pt Pre-DMD initializer, trained toward a Wan2.1-T2V-14B teacher
Distilled (1.3B teacher) distillalign_distill_1p3b_teacher.pt Final joint-distilled generator, Wan2.1-T2V-1.3B DMD teacher
Distilled (14B teacher) distillalign_distill_14b_teacher.pt Final joint-distilled generator, Wan2.1-T2V-14B DMD teacher

Each checkpoint stores a plain {"generator": state_dict} and loads directly with the inference and evaluation code in the supplementary material.

Teacher Reference Caches

Ready-made teacher reference features for the teacher-normalized distribution evaluation. Pass the .npz file to --teacher-features and only the student side needs to be generated.

Cache Download Description
Wan2.1-1.3B teacher reference features wan2.1_t2v_1.3b_reference_vjepa2.npz 256 x 2560 V-JEPA2 features, ready for --teacher-features
Wan2.1-14B teacher reference features wan2.1_t2v_14b_reference_vjepa2.npz 256 x 2560 V-JEPA2 features, ready for --teacher-features

See teacher_caches/README.md for the exact sampling and extraction protocol.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support