--- title: SGF+ Video Generation emoji: 🎬 colorFrom: purple colorTo: yellow sdk: gradio sdk_version: 6.29.1 app_file: app.py python_version: "3.10" startup_duration_timeout: 1h short_description: Few-step autoregressive text-to-video with SGF+ dual experts --- # SGF+: Decoupling Gradient Flows for Autoregressive Video Generation Interactive demo of **SGF+** (chunkwise Self Gradient Forcing Plus), a few-step autoregressive text-to-video model built on Wan2.1-T2V-1.3B. Each latent block (3 latents = 12 pixel frames) is denoised in only **4 steps** by the generation expert, then written back to the KV/cross-attention caches through a dedicated **memory expert** — the paper's core contribution — before the block is decoded and streamed to your browser as MPEG-TS chunks. - 📄 [Paper](https://huggingface.co/papers/2610.10429) - 💻 [Code (Apache-2.0)](https://github.com/Zihan-Su/Self_Gradient_Forcing_Plus) - 🤗 [Model weights](https://huggingface.co/ZihanSu/Self_Gradient_Forcing_Plus) ## Usage Type a prompt (or pick an example) and press **Generate video**. Frames start streaming as soon as the first block is ready; 7 blocks produce an 81-frame 480×832 clip at 16 fps (~5 s of video). Use the **Blocks** slider for shorter clips and the seed for reproducible generations. The inference loop follows the authors' `inference.py` (chunkwise config `sgf_plus_chunkwise.yaml`, EMA generator weights, warped 4-step denoising schedule, per-block memory-expert KV refresh) — capped to the 21-latent clip length the model was trained on, rather than the paper's long-horizon rollout. ## Credits SGF+ code and prompts are from the [Self_Gradient_Forcing_Plus](https://github.com/Zihan-Su/Self_Gradient_Forcing_Plus) repository (Apache-2.0). Base model: [Wan-AI/Wan2.1-T2V-1.3B](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B).