multimodalart's picture
multimodalart HF Staff
Upload folder using huggingface_hub
1ed0eb8 verified
|
Raw History Blame Contribute Delete
1.85 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade
metadata
title: SGF+ Video Generation
emoji: 🎬
colorFrom: purple
colorTo: yellow
sdk: gradio
sdk_version: 6.29.1
app_file: app.py
python_version: '3.10'
startup_duration_timeout: 1h
short_description: Few-step autoregressive text-to-video with SGF+ dual experts

SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

Interactive demo of SGF+ (chunkwise Self Gradient Forcing Plus), a few-step autoregressive text-to-video model built on Wan2.1-T2V-1.3B. Each latent block (3 latents = 12 pixel frames) is denoised in only 4 steps by the generation expert, then written back to the KV/cross-attention caches through a dedicated memory expert β€” the paper's core contribution β€” before the block is decoded and streamed to your browser as MPEG-TS chunks.

Usage

Type a prompt (or pick an example) and press Generate video. Frames start streaming as soon as the first block is ready; 7 blocks produce an 81-frame 480Γ—832 clip at 16 fps (~5 s of video). Use the Blocks slider for shorter clips and the seed for reproducible generations.

The inference loop follows the authors' inference.py (chunkwise config sgf_plus_chunkwise.yaml, EMA generator weights, warped 4-step denoising schedule, per-block memory-expert KV refresh) β€” capped to the 21-latent clip length the model was trained on, rather than the paper's long-horizon rollout.

Credits

SGF+ code and prompts are from the Self_Gradient_Forcing_Plus repository (Apache-2.0). Base model: Wan-AI/Wan2.1-T2V-1.3B.