Papers
arxiv:2609.38146

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

Published on Sep 29
· Submitted by
Shengxiang Ji
on Sep 30
Authors:
,
,
,
,
,
,
,
,
,
,
,

Abstract

We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.

Community

Paper submitter

TL;DR

  • Joint Camera and Future-Layout Control: LIFT is a unified video generation framework that enables users to
    control both camera motion and the semantic-spatial composition of newly revealed regions using only a
    last-frame layout.
  • Dual-mode OPSD Training: We use a dense spatiotemporal layout teacher to train a shared student in both the
    single-condition mode (conditioned only on the camera trajectory) and the dual-condition mode (conditioned on
    both the camera trajectory and the last-frame layout).
  • LIFT-Vista: We curate a dataset with large viewpoint changes and joint camera and layout annotations.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.38146
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.38146 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.