RE-SepFormer trained on Libri2Mix dataset

This model custom-trained Resource-Efficient Separation Transformer (RE-SepFormer) was designed to separate overlapping speech from audio mixtures into two separate sources.

Model Details

  • Developed by: Jahmori Richardson
  • Model type: Audio Source Separation (Deep Learning)
  • Language(s) (NLP): English
  • License: Apache 2.0

Model Sources [optional]

Uses

Isolating individual speakers from mixed audio files containing overlapping speech.

Downstream Use [optional]

[More Information Needed]

Out-of-Scope Use

It is not intended for non-speech audio separation or audio with heavy music/complex noise.

Bias, Risks, and Limitations

Model performance may vary with audio inputs that contain high amounts of background, reverberation, or more than required speakers that the model was designed for.

How to Get Started with the Model

Use the code below to get started with the model.

[More Information Needed]

Training Details

Training Data

[More Information Needed]

Training Procedure

Training Hyperparameters

  • Training regime: fp32 mixed precision
Downloads last month
137
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Jahmori-R/resepformer-librimix2spk 1

Paper for Jahmori-R/resepformer-librimix2spk