File size: 2,871 Bytes
6521f15
 
38d444b
 
 
6521f15
38d444b
 
 
 
 
824a59f
38d444b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
824a59f
 
 
 
 
49e254f
 
 
 
 
824a59f
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
---
license: apache-2.0
pipeline_tag: image-to-video
tags:
- video-generation
---

# LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

<p align="center">
  <a href="https://jsxzs.github.io/LIFT"><img src="https://img.shields.io/badge/Project%20Page-LIFT-blue?logo=googlechrome&logoColor=white" alt="Project Page"></a>
  <a href="https://arxiv.org/abs/2609.38146"><img src="https://img.shields.io/badge/arXiv-2609.38146-b31b1b?logo=arxiv&logoColor=white" alt="arXiv"></a>
  <a href="https://github.com/jsxzs/LIFT"><img src="https://img.shields.io/badge/GitHub-Code-181717?logo=github&logoColor=white" alt="Code"></a>
  <a href="https://huggingface.co/Overdog/LIFT"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-ffcc00" alt="Hugging Face Model"></a>
  <a href="https://huggingface.co/datasets/Overdog/LIFT-Vista"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-ffcc00" alt="Hugging Face Dataset"></a>
</p>

<p align="center">
  <img src="assets/teaser.png" width="720" alt="LIFT teaser">
</p>
<p align="center">
  <em>Given a first frame, users can navigate from the first-frame view along a desired camera path and specify layouts using bounding boxes with local text prompts in the final frame. Then, <b>LIFT</b> generates the intended shot that transitions from the input image to the user-defined last-frame layout following the prescribed camera trajectory.</em>
</p>

We introduce **LIFT**, a unified image-to-video generation framework that complements camera control with
**L**ayout-**I**n-**F**u**T**ure control, enabling users to specify what should appear in a future view and where it
should appear.

## Model Checkpoints

Our models are built on [Wan2.1-Fun-V1.1-1.3B-Control-Camera](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera).

| Model | Description |
|---|---|
| `LIFT/transformer/` | **Last-frame layout student** trained by dual-mode on-policy self-distillation from the dense-layout teacher. |
| `LIFT_dense_layout_teacher/transformer/` | **Dense-layout teacher** fine-tuned with dense per-frame layout |

## Download

```bash
pip install -U "huggingface_hub[cli]"
# Last-frame layout student (dual-mode OPSD)
hf download Overdog/LIFT --include "LIFT/*" --local-dir models
# Dense-layout teacher
hf download Overdog/LIFT --include "LIFT_dense_layout_teacher/*" --local-dir models
```

## Citation

```bibtex
@article{ji2026lift,
  title={LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation},
  author={Ji, Shengxiang and Wang, Boyang and Xu, Haiyang and Li, Bingnan and Mao, Yucheng and Chen, Zeyuan and Shan, Xiaojun and Zhang, Xiang and Hua, Gang and Xie, Jianwen and Cheng, Zezhou and Tu, Zhuowen},
  journal={arXiv preprint arXiv:2609.38146},
  year={2026}
}
```