--- license: other pipeline_tag: text-to-video tags: - video-generation - text-to-video - diffusion - distribution-matching - distillation - wan - arxiv:2604.03118 --- # 🧂 Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation [![arXiv](https://img.shields.io/badge/Arxiv-2604.03118-b31b1b)](https://arxiv.org/abs/2604.03118) [![Project Page](https://img.shields.io/badge/Project-Page-green)](https://xingtongge.github.io/Salt/) [![GitHub Repo stars](https://img.shields.io/github/stars/XingtongGe/Salt.svg?style=social&label=Star&maxAge=60)](https://github.com/XingtongGe/Salt) [![Hugging Face](https://img.shields.io/badge/🤗%20Hugging%20Face-Model-yellow)](https://huggingface.co/domiso/Salt) [Xingtong Ge](https://xingtongge.github.io/)1,2, [Yi Zhang](https://zhangyi-3.github.io/)2, Yushi Huang1, Dailan He2, Xiahong Wang2, Bingqi Ma2, Guanglu Song2, Yu Liu2, Jun Zhang1 1The Hong Kong University of Science and Technology, 2Vivix Group Limited European Conference on Computer Vision (**ECCV**), 2026 ## Abstract Distilling video generation models to extremely low inference budgets (e.g., 2-4 NFEs) is crucial for real-time deployment, yet remains challenging. Trajectory-style consistency distillation often becomes conservative under complex video dynamics, yielding over-smoothed appearance and weak motion. Distribution matching distillation (DMD) can recover sharp, mode-seeking samples, but its local training signals do not explicitly regularize how denoising updates compose across timesteps, making composed rollouts prone to drift. To overcome this challenge, we propose Self-Consistent Distribution Matching Distillation (SC-DMD), which explicitly regularizes the endpoint-consistent composition of consecutive denoising updates. For real-time autoregressive video generation, we further treat the KV cache as a quality-parameterized condition and propose cache-distribution-aware training. This training scheme applies SC-DMD over multi-step rollouts and introduces a cache-conditioned feature alignment objective that steers low-quality outputs toward high-quality references. Across extensive experiments on both non-autoregressive backbones (e.g., Wan 2.1) and autoregressive real-time paradigms (e.g., Self Forcing, Causal Forcing, and LongLive), Salt consistently improves low-NFE video generation quality while remaining compatible with diverse KV-cache memory mechanisms. ## Released Models ### Salt + Causal Forcing - `checkpoints/salt_cf.pt` - Supports 2-step and 4-step autoregressive generation. - Use the EMA generator for inference. ### Salt + LongLive - `checkpoints/salt_ll.pt` - Supports 4-step autoregressive generation with LongLive KV-cache memory. - Use the regular generator for inference. The training prompt collection used by the released recipes is also included at `prompts/vidprom_filtered_extended.txt`. ## Selected Results ### Text-to-video generation on VBench | Model | NFE | Total | Quality | Semantic | | --- | ---: | ---: | ---: | ---: | | Self Forcing | 4 | 84.20 | 84.74 | **82.05** | | Salt + Self Forcing | 4 | **84.47** | **85.27** | 81.28 | | LongLive | 4 | 84.40 | 85.12 | 81.53 | | Salt + LongLive | 4 | **84.93** | **85.41** | **83.00** | | Causal Forcing | 4 | 84.62 | 85.41 | 81.47 | | Salt + Causal Forcing | 4 | **85.08** | **85.96** | **81.59** | | Salt + Causal Forcing | 2 | **84.80** | **85.63** | **81.49** | ![Qualitative comparisons with Causal Forcing](https://raw.githubusercontent.com/XingtongGe/Salt/main/assets/vis_supp.png) ## Usage ### 1. Prepare the code and artifacts ```bash git clone https://github.com/XingtongGe/Salt.git cd Salt # Downloads checkpoints/ and prompts/ into the paths expected by the configs. hf download domiso/Salt --local-dir . ``` Follow the installation and Wan2.1 preparation instructions in the [GitHub repository](https://github.com/XingtongGe/Salt). ### 2. Salt + Causal Forcing ```bash python inference.py \ --config_path configs/inference/salt_causal_forcing.yaml \ --checkpoint_path checkpoints/salt_cf.pt \ --data_path prompts/example_prompts.txt \ --output_folder outputs/salt_cf \ --use_ema ``` ### 3. Salt + LongLive ```bash python inference.py \ --config_path configs/inference/salt_longlive.yaml \ --checkpoint_path checkpoints/salt_ll.pt \ --data_path prompts/example_prompts.txt \ --output_folder outputs/salt_ll ``` ## Training The released prompt collection contains one prompt per line and matches the default public recipes. Salt training does not require a video dataset. The code repository provides recipes for: - Self Forcing, Causal Forcing, and LongLive baselines; - mixed-step SC-DMD; and - mixed-step SC-DMD with TRD alignment. All canonical mixed-step recipes sample 8-, 4-, and 2-step trajectories with probabilities **0.4 / 0.4 / 0.2**. See the [configuration guide](https://github.com/XingtongGe/Salt/tree/main/configs) for the complete recipe matrix. ## License Please review the code repository's license and third-party notices before redistribution or commercial use. In particular, the LongLive backbone carries a file-level CC-BY-NC-SA-4.0 notice. Model weights may also be subject to their upstream backbone licenses. ## Citation If you find this work useful, please cite: ```bibtex @article{ge2026salt, title={Salt: Self-consistent distribution matching with cache-aware training for fast video generation}, author={Ge, Xingtong and Zhang, Yi and Huang, Yushi and He, Dailan and Wang, Xiahong and Ma, Bingqi and Song, Guanglu and Liu, Yu and Zhang, Jun}, journal={arXiv preprint arXiv:2604.03118}, year={2026} } ```