Image-to-Video
English
CausalWM / README.md
Maxceline's picture
Add arXiv citation for CausalWM
a1d46dc verified
|
Raw History Blame Contribute Delete
2.71 kB
---
language:
- en
license: other
license_name: ltx-2-community
license_link: LICENSE
pipeline_tag: image-to-video
---
# CausalWMv1
The current released checkpoint is tailored for **PAI-Bench-G**. Additional checkpoints will be released soon.
CausalWMv1 takes one RGB observation and a text instruction, then generates optical
flow, camera-frame XYZ pointmaps, and future RGB video through a stage-causal chain.
- Code: [AetherLabsAI/CausalWM](https://github.com/AetherLabsAI/CausalWM)
- Paper: [CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model](https://arxiv.org/pdf/2609.23184)
- Checkpoint: `CausalWMv1.safetensors`
## Files and requirements
The checkpoint contains the complete CoT Transformer state in BF16. It is not a
standalone pipeline: inference additionally needs `ltx-2.3-22b-dev.safetensors`
from [Lightricks/LTX-2.3](https://huggingface.co/Lightricks/LTX-2.3) and the
[Gemma-3-12B text encoder](https://huggingface.co/google/gemma-3-12b-it-qat-q4_0-unquantized).
See `model-manifest.json` for the file size, SHA-256, tensor count, and runtime
metadata. Verify the download using `sha256sum -c SHA256SUMS`.
## Inference
Follow the repository installation instructions for TI2V inference, then run:
```bash
python inference.py \
--image /path/to/first_frame.png \
--prompt "The robot arm picks up the red block and places it in the box." \
--checkpoint /path/to/CausalWMv1.safetensors \
--base-ckpt /path/to/ltx-2.3-22b-dev.safetensors \
--text-encoder-dir /path/to/gemma-3-12b \
--out-dir outputs/demo
```
Default settings are 121 frames, 640×480, 16 FPS, four steps per stage, guidance 1.0,
and seed 42. The input frame is resized without cropping. Pointmaps use a relative
camera-frame scale normalized by the first frame's median depth, not metres.
Generated motion and geometry are predictions rather than measured ground truth.
## License and attribution
See `LICENSE` and `NOTICE` for the LTX-2 license and attribution carried with this
release. Base model and text encoder downloads remain subject to their respective
terms. This package contains inference weights, not optimizer or training state.
## Citation
If CausalWM is useful for your research, please cite our paper:
```bibtex
@misc{xu2026causalwmcausalchainofthoughtreasoning,
title={CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model},
author={Ziming Xu and Shuang Liang and Ruobing Han and Ziqiao Xi and Mingxing Rao and Kun Zhou and Zijun Zhang and Yuchen Yan and Yufan Wei and Junbo Huang and Yifei Shao and Fang Nan and Biwei Huang},
year={2026},
eprint={2609.23184},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.23184},
}
```