LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

Project Page Code Hugging Face Model Hugging Face Dataset

LIFT teaser

Given a first frame, users can navigate from the first-frame view along a desired camera path and specify layouts using bounding boxes with local text prompts in the final frame. Then, LIFT generates the intended shot that transitions from the input image to the user-defined last-frame layout following the prescribed camera trajectory.

We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear.

Model Checkpoints

Our models are built on Wan2.1-Fun-V1.1-1.3B-Control-Camera.

Model Description
LIFT/transformer/ Last-frame layout student trained by dual-mode on-policy self-distillation from the dense-layout teacher.
LIFT_dense_layout_teacher/transformer/ Dense-layout teacher fine-tuned with dense per-frame layout

Download

pip install -U "huggingface_hub[cli]"
# Last-frame layout student (dual-mode OPSD)
hf download Overdog/LIFT --include "LIFT/*" --local-dir models
# Dense-layout teacher
hf download Overdog/LIFT --include "LIFT_dense_layout_teacher/*" --local-dir models
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support