SepGen generation checkpoint (gen-3k)
SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model
Aviad Dahan1, Rajaei Khatib1, Yonatan Bitton2, Idan Szpektor2, Lior Wolf1, Raja Giryes1
1Tel Aviv University, 2Google
Paper · Project page · Code · Other checkpoint: sep-12k (separation)
SepGen extends LTX-2.5 to emit the video, the audio-mix, and one waveform per captioned source in a single sampling run. The same weights generate (scene caption and per-source captions → video, audio-mix and stems) and separate (observed video and audio-mix → stems), selected by the noise level of the audio-mix.
gen-3k is the generation checkpoint: sep-12k continued for 3,000 steps in which 70% of samples present the audio-mix at a noise level no higher than the stems'. From a scene caption and one caption per source it generates the video, the audio-mix and one waveform per source; it also separates, at parity with sep-12k.
Files
| File | Content |
|---|---|
lora_weights.safetensors |
the LoRA adapter (sha256 a72b3c662ce8d6b6e29013bbedb797d48ee1c90e8b7d219af4a27329eb58515b) |
config.json |
adapter and training summary |
train_config.yaml |
the full training config |
The adapter targets the audio stream of the LTX-2.5 22B dev transformer: rank and alpha 128, 480
modules, 327M parameters, bf16, ComfyUI key format (diffusion_model.). It is trained with
attention gates (block-diagonal caption routing, the protected audio-mix span, video reading audio
from the audio-mix only, Cross-Stem Attention Guidance) that must also be installed at inference,
so run it with the SepGen code, not as a plain LoRA.
Usage
git clone https://github.com/AviadDahan/SepGen && cd SepGen
bash setup.sh && source activate.sh && bash download_weights.sh
python generate.py examples/generation_prompts.json --out-dir outputs/gen # uses gen-3k
python separate.py --batch-manifest examples/separation_manifest.json --checkpoint gen-3k --out-dir outputs/sep
--checkpoint gen-3k downloads this repo on first use.
License
These weights are a Derivative of LTX-2.x and are distributed under the LTX-2.x Community License Agreement, including its use-based restrictions (Attachment A) and the LTX Acceptable Use Policy.
Citation
@article{dahan2026sepgen,
title = {SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model},
author = {Dahan, Aviad and Khatib, Rajaei and Bitton, Yonatan and Szpektor, Idan and Wolf, Lior and Giryes, Raja},
journal = {arXiv preprint arXiv:2610.11361},
year = {2026}
}
- Downloads last month
- 2
Model tree for AviadDahan/SepGen-Generation
Base model
Lightricks/LTX-2.5