|
Download README.md from wq2012/vanilla_seq2seq_multi_interfering: direct link, hf CLI and curl.
- Browser
- Download file 6.96 kB
-
https://huggingface.co/wq2012/vanilla_seq2seq_multi_interfering/resolve/main/README.md
- Command line
-
hf download hf://wq2012/vanilla_seq2seq_multi_interfering/README.md
-
curl -L -o README.md https://huggingface.co/wq2012/vanilla_seq2seq_multi_interfering/resolve/main/README.md
6.96 kB
| license: apache-2.0 | |
| library_name: lingvo | |
| tags: | |
| - audio | |
| - speech-enhancement | |
| - echo-cancellation | |
| - textual-echo-cancellation | |
| - tensorflow | |
| - tflite | |
| datasets: | |
| - libritts | |
| - vctk | |
| metrics: | |
| - wer | |
| - mcd | |
| arxiv: 2008.06006 | |
| # Vanilla-Seq2seq Baseline (No Side Input) β Multiple Interfering Voices (`NoSideInputMultiInterfering`) | |
| Single-source sequence-to-sequence speech enhancement baseline (`Vanilla-Seq2seq` / `NoSideInput`) trained on 24 kHz **LibriTTS** user speech mixed at 0 dB SNR with reverberant **CSTR VCTK** without any side input. | |
| - **GitHub Repository**: [https://github.com/wq2012/tec](https://github.com/wq2012/tec) | |
| - **PyPI Package**: [`textual-echo-cancellation`](https://pypi.org/project/textual-echo-cancellation/) | |
| - **Paper**: [Textual Echo Cancellation (IEEE SLT 2021, arXiv:2008.06006v4)](https://arxiv.org/pdf/2008.06006) | |
| - **Audio Demo Page**: [https://google.github.io/speaker-id/publications/TEC/](https://google.github.io/speaker-id/publications/TEC/) | |
| > **Open-Source Reproduction Notice**: This pretrained model was trained using the standalone open-source reproduction library ([`wq2012/tec`](https://github.com/wq2012/tec)) built on [`lingvo`](https://github.com/tensorflow/lingvo) and `tensorflow`. It does not use Google's proprietary internal codebase or internal training data infrastructure. | |
| --- | |
| ## Model Performance | |
| ### 1. Open-Source Reproduction Evaluation (`NoSideInputMultiInterfering`) | |
| Evaluated on 24 kHz **LibriTTS** (`test-clean`, `test-other`) mixed at 0 dB SNR with synthetic room impulse responses (RT60 = 0.25 s) under the **Multiple interfering voices (LibriTTS + VCTK)** condition, using `Qwen3-ASR-0.6B-F16` via `audio.cpp` for ASR WER scoring and 13-MFCC Dynamic Time Warping for MCD: | |
| | Metric | `test-clean` | `test-other` |Paper Reference (`Vanilla-Seq2seq`) | | |
| | :--- | :---: | :---: | :---: | | |
| | **WER (%) β** | **31.16%** (62/199) | **42.39%** (103/243) | 19.7% / 38.7% | | |
| | **MCD (dB) β** | **7.93 dB** | **8.72 dB** | 7.53 dB / 8.87 dB | | |
| | **Side Input Size (KB) β** | **0.000 KB** | **0.000 KB** | 0 KB | | |
| | **Computational Complexity β** | **6.32 GFLOPS** | **6.32 GFLOPS** | 6.32 GFLOPS | | |
| ### 2. Full Comparison Across All Pretrained Models & Baselines | |
| | Condition | Method | Hugging Face Model | WER (%) test-clean β | WER (%) test-other β | MCD (dB) test-clean β | MCD (dB) test-other β | Side Input test-clean (KB) β | | |
| | :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: | | |
| | **Single interfering voice** *(LibriTTS + LJSpeech)* | `GroundTruth` | β | 3.52 | 6.78 | 0.00 | 0.00 | 0.000 | | |
| | | `MicrophoneSignal` | β | 90.45 | 114.12 | 12.86 | 14.61 | 0.000 | | |
| | | `NlmsAec` (AEC-NLMS) | β | 88.44 | 107.34 | 12.80 | 14.48 | 243.465 | | |
| | | `NoSideInputSingleInterfering` | [`wq2012/vanilla_seq2seq_single_interfering`](https://huggingface.co/wq2012/vanilla_seq2seq_single_interfering) | 45.23 | 91.53 | 9.58 | 11.34 | 0.000 | | |
| | | `AecSingleInterfering` | [`wq2012/aec_single_interfering`](https://huggingface.co/wq2012/aec_single_interfering) | 12.06 | 23.16 | 8.85 | 9.86 | 243.465 | | |
| | | **`TecSingleInterfering`** | **[`wq2012/tec_single_interfering`](https://huggingface.co/wq2012/tec_single_interfering)** | **21.61** | **46.89** | **8.24** | **9.28** | **0.076** | | |
| | **Multiple interfering voices** *(LibriTTS + VCTK)* | `GroundTruth` | β | 5.03 | 7.82 | 0.00 | 0.00 | 0.000 | | |
| | | `MicrophoneSignal` | β | 34.17 | 48.97 | 7.67 | 7.70 | 0.000 | | |
| | | `NlmsAec` (AEC-NLMS) | β | 28.64 | 34.98 | 7.92 | 8.40 | 186.922 | | |
| | | `NoSideInputMultiInterfering` | [`wq2012/vanilla_seq2seq_multi_interfering`](https://huggingface.co/wq2012/vanilla_seq2seq_multi_interfering) | 31.16 | 42.39 | 7.93 | 8.72 | 0.000 | | |
| | | `AecMultiInterfering` | [`wq2012/aec_multi_interfering`](https://huggingface.co/wq2012/aec_multi_interfering) | 8.54 | 22.22 | 7.80 | 7.88 | 186.922 | | |
| | | **`TecMultiInterfering`** | **[`wq2012/tec_multi_interfering`](https://huggingface.co/wq2012/tec_multi_interfering)** | **26.63** | **45.27** | **7.96** | **8.40** | **0.037** | | |
| --- | |
| ## Files in This Repository | |
| - `best.ckpt.data-00000-of-00001`, `best.ckpt.index`, `best.ckpt.meta`, `checkpoint`: TensorFlow / Lingvo checkpoint for `NoSideInputMultiInterfering`. | |
| - `model.tflite`: Dynamic-range quantized TensorFlow Lite (`.tflite`) FlatBuffer model for on-device inference. | |
| - `evaluation_metrics.json`: Verified evaluation results (`test-clean` and `test-other`) for `NoSideInputMultiInterfering`. | |
| --- | |
| ## How to Use | |
| ### 1. Install `textual-echo-cancellation` | |
| ```bash | |
| pip3 install textual-echo-cancellation huggingface_hub | |
| ``` | |
| ### 2. Download the Model from Hugging Face | |
| ```python | |
| from huggingface_hub import snapshot_download | |
| model_dir = snapshot_download(repo_id="wq2012/vanilla_seq2seq_multi_interfering") | |
| print("Downloaded model to:", model_dir) | |
| ``` | |
| ### 3. Run Inference via CLI (`scripts/inference.py`) | |
| ```bash | |
| python3 -m scripts.inference \ | |
| --model NoSideInputMultiInterfering \ | |
| --checkpoint_path "${MODEL_DIR}/best.ckpt" \ | |
| --mixed_wav /path/to/mixed_input.wav \ | |
| # No side input required for Vanilla-Seq2seq \ | |
| --output_wav /tmp/enhanced_clean.wav | |
| ``` | |
| ### 4. Run Inference via Python API | |
| ```python | |
| import os | |
| from huggingface_hub import snapshot_download | |
| from tec import inference | |
| model_dir = snapshot_download(repo_id="wq2012/vanilla_seq2seq_multi_interfering") | |
| ckpt_path = os.path.join(model_dir, "best.ckpt") | |
| result = inference.run_inference_on_wav( | |
| model_name="NoSideInputMultiInterfering", | |
| mixed_wav_path="/path/to/mixed_input.wav", | |
| interfering_text="currently in mountain view it is 72 degrees", | |
| checkpoint_path=ckpt_path, | |
| output_wav_path="/tmp/enhanced_clean.wav", | |
| ) | |
| print("Enhanced log-Mel spectrogram shape:", result["predicted_mel"].shape) | |
| ``` | |
| ### 5. On-Device Inference with Quantized TFLite (`model.tflite`) | |
| ```python | |
| import os | |
| import numpy as np | |
| import tensorflow as tf | |
| from huggingface_hub import snapshot_download | |
| model_dir = snapshot_download(repo_id="wq2012/vanilla_seq2seq_multi_interfering") | |
| tflite_path = os.path.join(model_dir, "model.tflite") | |
| interpreter = tf.lite.Interpreter(model_path=tflite_path) | |
| interpreter.allocate_tensors() | |
| for detail in interpreter.get_input_details(): | |
| interpreter.set_tensor( | |
| detail["index"], np.zeros(detail["shape"], dtype=detail["dtype"]) | |
| ) | |
| interpreter.invoke() | |
| output_details = interpreter.get_output_details() | |
| enhanced_mel = interpreter.get_tensor(output_details[0]["index"]) | |
| print("TFLite predicted log-Mel shape:", enhanced_mel.shape) | |
| ``` | |
| --- | |
| ## Citation | |
| If you use this model or the `textual-echo-cancellation` library in your research, please cite the original paper: | |
| ```bibtex | |
| @inproceedings{ding2021textual, | |
| title={Textual Echo Cancellation}, | |
| author={Ding, Shaojin and Jia, Ye and Hu, Ke and Wang, Quan}, | |
| booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)}, | |
| pages={666--673}, | |
| year={2021}, | |
| organization={IEEE} | |
| } | |
| ``` | |