SepGen separation checkpoint (sep-12k)
SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model
Aviad Dahan1, Rajaei Khatib1, Yonatan Bitton2, Idan Szpektor2, Lior Wolf1, Raja Giryes1
1Tel Aviv University, 2Google
Paper · Project page · Code · Other checkpoint: gen-3k (generation)
SepGen extends LTX-2.5 to emit the video, the audio-mix, and one waveform per captioned source in a single sampling run. The same weights generate (scene caption and per-source captions → video, audio-mix and stems) and separate (observed video and audio-mix → stems), selected by the noise level of the audio-mix.
sep-12k is the separation checkpoint: 12,000 training steps in which the audio-mix is always presented clean and the loss is on the two stems. Given a video, its audio-mix and one caption per source, it returns one waveform per source.
Files
| File | Content |
|---|---|
lora_weights.safetensors |
the LoRA adapter (sha256 fc0987e972c0e92556348ea99b018c9a7c3a95691c4bd46a7d549b1fcd4edd39) |
config.json |
adapter and training summary |
train_config.yaml |
the full training config |
The adapter targets the audio stream of the LTX-2.5 22B dev transformer: rank and alpha 128, 480
modules, 327M parameters, bf16, ComfyUI key format (diffusion_model.). It is trained with
attention gates (block-diagonal caption routing, the protected audio-mix span, video reading audio
from the audio-mix only, Cross-Stem Attention Guidance) that must also be installed at inference,
so run it with the SepGen code, not as a plain LoRA.
Usage
git clone https://github.com/AviadDahan/SepGen && cd SepGen
bash setup.sh && source activate.sh && bash download_weights.sh
python separate.py --batch-manifest examples/separation_manifest.json --checkpoint sep-12k --out-dir outputs/sep
--checkpoint sep-12k downloads this repo on first use.
License
These weights are a Derivative of LTX-2.x and are distributed under the LTX-2.x Community License Agreement, including its use-based restrictions (Attachment A) and the LTX Acceptable Use Policy.
Citation
@article{dahan2026sepgen,
title = {SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model},
author = {Dahan, Aviad and Khatib, Rajaei and Bitton, Yonatan and Szpektor, Idan and Wolf, Lior and Giryes, Raja},
journal = {arXiv preprint arXiv:2610.11361},
year = {2026}
}
- Downloads last month
- 2
Model tree for AviadDahan/SepGen-Separation
Base model
Lightricks/LTX-2.5