## Table of Contents
- [Quick Start](#quick-start)
- [Installation](#installation)
- [Inference: 1/2/4-step T2V & I2V](#cli-inference)
- [Long Video](#minute-level-long-video-generation)
- [Training Pipeline](#training) (Causal Foricng / Causal Forcing++)
- [Stage 1: AR Diffusion](#training)
- [Stage 2: Causal ODE (Causal Forcing)](#training) / [🔥Causal CD (Causal Forcing++)](#-stage-2-option-b-causal-consistency-distillation-initialization-causal-forcing)
- [Stage 3: Asymmetric DMD](#stage-3-dmd)
| Models | Checkpoints | Description|
|--------------------|---------------------------------------------------------------------------------------------------------------------------------------------|-------|
| Chunk-wise 4-step | 🤗 [Huggingface](https://huggingface.co/zhuhz22/Causal-Forcing/blob/main/chunkwise/causal_forcing.pt) | SOTA AR model on Wan1.3B outperforming Self Forcing|
| Frame-wise 4-step | 🤗 [Huggingface](https://huggingface.co/zhuhz22/Causal-Forcing/blob/main/framewise/causal_forcing.pt) | Rich dynamics and high quality |
| Frame-wise 2-step🔥 | 🤗 [Huggingface](https://huggingface.co/zhuhz22/Causal-Forcing/blob/main/causal-forcing%2B%2B/framewise-2step.pt) | The first frame-wise 2-step model **even better than 4-step**|
| Frame-wise 1-step | 🤗 [Huggingface](https://huggingface.co/zhuhz22/Causal-Forcing/blob/main/causal-forcing%2B%2B/framewise-1step.pt) | Extremely low latency |
| Long Video Generator 🔥 | 🤗 [Huggingface](https://huggingface.co/zhuhz22/Causal-Forcing/blob/main/chunkwise/longvideo.pt) | Minute-long AR video generator |
| HY1.5-TI2V (8B) 🔥 | 🤗 [Huggingface](https://huggingface.co/MIN-Lab/minWM) | Refer to [this repo](https://github.com/shengshu-ai/minWM). |
| Action-conditioned WM | 🤗 [Huggingface](https://huggingface.co/MIN-Lab/minWM) | Refer to [this repo](https://github.com/shengshu-ai/minWM). |
-----
https://github.com/user-attachments/assets/310f0cfa-e1bb-496d-8941-87f77b3271c0
## 🔥 News
- **2026.7.23**: [Self Gradient Forcing](https://github.com/zhuang2002/Self_Gradient_Forcing) is built on Causal Forcing initialization.
- **2026.7.20**: Happy to see that the recent SOTA video world models [DreamX-World 1.0](https://arxiv.org/pdf/2606.16993) and [Matrix-Game 3.5](https://matrix-game-v3-5.github.io/paper/Matrix-Game-3.5.pdf) are built on Causal Forcing!
- **2026.5.17**: We release Causal Forcing for the HY1.5-TI2V-8B model! Refer to [this repo](https://github.com/shengshu-ai/minWM) for the details. This model explicitly supports I2V.
- **2026.5.15**: We release [Causal Forcing++](https://arxiv.org/abs/2605.15141), supporting Casual Consistency Distillation for few-step initialization, and open-source **the first frame-wise 2-step AR model** comparable to chunk-wise 4-step models!
- **2026.5.10**: Thanks to @[AshadowZ](https://github.com/AshadowZ), now our chunk-wise ODE data curation is **3x faster**!
- **2026.4.16**: **Optimize the Stage 2 🔥consistency distillation🔥 infrastructure for 3× faster training, let's try it now!** We have also released the ckpt.
- **2026.3.15** : [Rolling Sink](https://github.com/haodong2000/RollingSink), [Infinity-RoPE](https://github.com/yesiltepe-hidir/infinity-rope) and [Deep Forcing](https://cvlab-kaist.github.io/DeepForcing/) adopt Causal Forcing as one of the base models!
- **2026.2.28** : Add [FAQ section](#faq--blog) regarding hot topics, specifically which is the better Initialization between AR diffusion and causal ODE distillation.
- **2026.2.11** : We now support **I2V** generation! Feel free to try it [here](#new-i2v)!
- **2026.2.7** : Causal Forcing now supports [Rolling Forcing](https://github.com/TencentARC/RollingForcing), enabling minute-level long video generation!
- **2026.2.5** : Release causal consistency distillation (Preview) as substitute for ODE distillation, **free of generating ODE paired data**!
- **2026.2.2** : The [paper](https://arxiv.org/abs/2602.02214), [project page](https://thu-ml.github.io/CausalForcing.github.io/), and code are released.
## Quick Start
> The inference environment is identical to Self Forcing.
**NOTE**: Similar to CausVid/Self Forcing, Causal Forcing does not natively support videos longer than 81 frames. As a base training method, it is orthogonal to techniques like Longlive/Rolling Forcing. To use Causal Forcing as a long video baseline, see [this extension](#minute-level-long-video-generation). **Directly using the 5-second trained Causal Forcing model as a baseline for long video generation is extremely unfair**.
### Installation
```bash
conda create -n causal_forcing python=3.10 -y
conda activate causal_forcing
pip install -r requirements.txt
pip install git+https://github.com/openai/CLIP.git
pip install flash-attn --no-build-isolation
python setup.py develop
```
### Download Checkpoints
```bash
hf download Wan-AI/Wan2.1-T2V-1.3B --local-dir wan_models/Wan2.1-T2V-1.3B
hf download Wan-AI/Wan2.1-T2V-14B --local-dir wan_models/Wan2.1-T2V-14B
# Causal Forcing
hf download zhuhz22/Causal-Forcing chunkwise/causal_forcing.pt --local-dir checkpoints
hf download zhuhz22/Causal-Forcing framewise/causal_forcing.pt --local-dir checkpoints
# Causal Forcing++
hf download zhuhz22/Causal-Forcing causal-forcing++/framewise-2step.pt --local-dir checkpoints
hf download zhuhz22/Causal-Forcing causal-forcing++/framewise-1step.pt --local-dir checkpoints
```
### CLI Inference
#### T2V
Chunk-wise model:
```bash
python inference.py \
--config_path configs/causal_forcing_dmd_chunkwise.yaml \
--output_folder output/chunkwise \
--checkpoint_path checkpoints/chunkwise/causal_forcing.pt \
--data_path prompts/demos.txt
```
Frame-wise model:
```bash
# =============== Causal Forcing++ ================
# 2-step Causal Forcing++
python inference.py \
--config_path configs/causal_forcing_dmd_framewise_2step.yaml \
--output_folder output/framewise_2step_cf++ \
--checkpoint_path checkpoints/causal-forcing++/framewise-2step.pt \
--data_path prompts/demos.txt \
--use_ema
# 1-step Causal Forcing++
python inference.py \
--config_path configs/causal_forcing_dmd_framewise_1step.yaml \
--output_folder output/framewise_1step_cf++ \
--checkpoint_path checkpoints/causal-forcing++/framewise-1step.pt \
--data_path prompts/demos.txt \
--use_ema
# =============== Causal Forcing ================
# 4-step Causal Forcing
python inference.py \
--config_path configs/causal_forcing_dmd_framewise.yaml \
--output_folder output/framewise \
--checkpoint_path checkpoints/framewise/causal_forcing.pt \
--data_path prompts/demos.txt \
--use_ema
```
#### I2V
> Our frame-wise setting natively supports I2V. You simply need to set the first latent initial frame as your conditional image.
```bash
python inference.py \
--config_path configs/causal_forcing_dmd_framewise.yaml \
--output_folder output/framewise \
--checkpoint_path checkpoints/framewise/causal_forcing.pt \
--data_path prompts/i2v \
--i2v \
--use_ema
```
### Minute-level Long Video Generation
Built on [Rolling Forcing](https://github.com/TencentARC/RollingForcing), we implemented minute-level long video generation. See [here](./long_video) for the detail.
[Infinity-RoPE](https://github.com/yesiltepe-hidir/infinity-rope), [Deep Forcing](https://cvlab-kaist.github.io/DeepForcing/) and [Rolling Sink](https://github.com/haodong2000/RollingSink) also adopt Causal Forcing as one of their base models, enabling interactive (prompt-switchable) long video generation at the minute scale. You can also try them out at their repos.
## Training