SepGen separation checkpoint (sep-12k)

SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model

Aviad Dahan1, Rajaei Khatib1, Yonatan Bitton2, Idan Szpektor2, Lior Wolf1, Raja Giryes1

1Tel Aviv University, 2Google

Paper · Project page · Code · Other checkpoint: gen-3k (generation)

SepGen extends LTX-2.5 to emit the video, the audio-mix, and one waveform per captioned source in a single sampling run. The same weights generate (scene caption and per-source captions → video, audio-mix and stems) and separate (observed video and audio-mix → stems), selected by the noise level of the audio-mix.

sep-12k is the separation checkpoint: 12,000 training steps in which the audio-mix is always presented clean and the loss is on the two stems. Given a video, its audio-mix and one caption per source, it returns one waveform per source.

Files

File Content
lora_weights.safetensors the LoRA adapter (sha256 fc0987e972c0e92556348ea99b018c9a7c3a95691c4bd46a7d549b1fcd4edd39)
config.json adapter and training summary
train_config.yaml the full training config

The adapter targets the audio stream of the LTX-2.5 22B dev transformer: rank and alpha 128, 480 modules, 327M parameters, bf16, ComfyUI key format (diffusion_model.). It is trained with attention gates (block-diagonal caption routing, the protected audio-mix span, video reading audio from the audio-mix only, Cross-Stem Attention Guidance) that must also be installed at inference, so run it with the SepGen code, not as a plain LoRA.

Usage

git clone https://github.com/AviadDahan/SepGen && cd SepGen
bash setup.sh && source activate.sh && bash download_weights.sh
python separate.py --batch-manifest examples/separation_manifest.json --checkpoint sep-12k --out-dir outputs/sep

--checkpoint sep-12k downloads this repo on first use.

License

These weights are a Derivative of LTX-2.x and are distributed under the LTX-2.x Community License Agreement, including its use-based restrictions (Attachment A) and the LTX Acceptable Use Policy.

Citation

@article{dahan2026sepgen,
  title   = {SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model},
  author  = {Dahan, Aviad and Khatib, Rajaei and Bitton, Yonatan and Szpektor, Idan and Wolf, Lior and Giryes, Raja},
  journal = {arXiv preprint arXiv:2610.11361},
  year    = {2026}
}
Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AviadDahan/SepGen-Separation

Adapter
(34)
this model

Collection including AviadDahan/SepGen-Separation

Paper for AviadDahan/SepGen-Separation