--- license: apache-2.0 library_name: lingvo tags: - audio - speech-enhancement - echo-cancellation - textual-echo-cancellation - tensorflow - tflite datasets: - libritts - ljspeech metrics: - wer - mcd arxiv: 2008.06006 --- # AEC-Seq2seq Neural Baseline — Single Interfering Voice (`AecSingleInterfering`) Multi-source attention sequence-to-sequence **Acoustic Echo Cancellation (`AEC-Seq2seq`)** neural baseline model trained on 24 kHz **LibriTTS** user speech mixed at 0 dB SNR with reverberant **LJ Speech** (single-speaker TTS interference). Uses dual audio encoders (`SpeechEncoderV1`) on both the microphone mixture and the full reference TTS playback audio (~240–310 KB side input). - **GitHub Repository**: [https://github.com/wq2012/tec](https://github.com/wq2012/tec) - **PyPI Package**: [`textual-echo-cancellation`](https://pypi.org/project/textual-echo-cancellation/) - **Paper**: [Textual Echo Cancellation (IEEE SLT 2021, arXiv:2008.06006v4)](https://arxiv.org/pdf/2008.06006) - **Audio Demo Page**: [https://google.github.io/speaker-id/publications/TEC/](https://google.github.io/speaker-id/publications/TEC/) > **Open-Source Reproduction Notice**: This pretrained model was trained using the standalone open-source reproduction library ([`wq2012/tec`](https://github.com/wq2012/tec)) built on [`lingvo`](https://github.com/tensorflow/lingvo) and `tensorflow`. It does not use Google's proprietary internal codebase or internal training data infrastructure. --- ## Model Performance ### 1. Open-Source Reproduction Evaluation (`AecSingleInterfering`) Evaluated on 24 kHz **LibriTTS** (`test-clean`, `test-other`) mixed at 0 dB SNR with synthetic room impulse responses (RT60 = 0.25 s) under the **Single interfering voice (LibriTTS + LJ Speech)** condition, using `Qwen3-ASR-0.6B-F16` via `audio.cpp` for ASR WER scoring and 13-MFCC Dynamic Time Warping for MCD: | Metric | `test-clean` | `test-other` |Paper Reference (`AEC-Seq2seq`) | | :--- | :---: | :---: | :---: | | **WER (%) ↓** | **12.06%** (24/199) | **23.16%** (41/177) | 8.30% / 24.3% | | **MCD (dB) ↓** | **8.85 dB** | **9.86 dB** | 6.38 dB / 7.07 dB | | **Side Input Size (KB) ↓** | **243.465 KB** | **209.085 KB** | 310 KB | | **Computational Complexity ↓** | **9.51 GFLOPS** | **9.51 GFLOPS** | 9.51 GFLOPS | ### 2. Full Comparison Across All Pretrained Models & Baselines | Condition | Method | Hugging Face Model | WER (%) test-clean ↓ | WER (%) test-other ↓ | MCD (dB) test-clean ↓ | MCD (dB) test-other ↓ | Side Input test-clean (KB) ↓ | | :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: | | **Single interfering voice** *(LibriTTS + LJSpeech)* | `GroundTruth` | — | 3.52 | 6.78 | 0.00 | 0.00 | 0.000 | | | `MicrophoneSignal` | — | 90.45 | 114.12 | 12.86 | 14.61 | 0.000 | | | `NlmsAec` (AEC-NLMS) | — | 88.44 | 107.34 | 12.80 | 14.48 | 243.465 | | | `NoSideInputSingleInterfering` | [`wq2012/vanilla_seq2seq_single_interfering`](https://huggingface.co/wq2012/vanilla_seq2seq_single_interfering) | 45.23 | 91.53 | 9.58 | 11.34 | 0.000 | | | `AecSingleInterfering` | [`wq2012/aec_single_interfering`](https://huggingface.co/wq2012/aec_single_interfering) | 12.06 | 23.16 | 8.85 | 9.86 | 243.465 | | | **`TecSingleInterfering`** | **[`wq2012/tec_single_interfering`](https://huggingface.co/wq2012/tec_single_interfering)** | **21.61** | **46.89** | **8.24** | **9.28** | **0.076** | | **Multiple interfering voices** *(LibriTTS + VCTK)* | `GroundTruth` | — | 5.03 | 7.82 | 0.00 | 0.00 | 0.000 | | | `MicrophoneSignal` | — | 34.17 | 48.97 | 7.67 | 7.70 | 0.000 | | | `NlmsAec` (AEC-NLMS) | — | 28.64 | 34.98 | 7.92 | 8.40 | 186.922 | | | `NoSideInputMultiInterfering` | [`wq2012/vanilla_seq2seq_multi_interfering`](https://huggingface.co/wq2012/vanilla_seq2seq_multi_interfering) | 31.16 | 42.39 | 7.93 | 8.72 | 0.000 | | | `AecMultiInterfering` | [`wq2012/aec_multi_interfering`](https://huggingface.co/wq2012/aec_multi_interfering) | 8.54 | 22.22 | 7.80 | 7.88 | 186.922 | | | **`TecMultiInterfering`** | **[`wq2012/tec_multi_interfering`](https://huggingface.co/wq2012/tec_multi_interfering)** | **26.63** | **45.27** | **7.96** | **8.40** | **0.037** | --- ## Files in This Repository - `best.ckpt.data-00000-of-00001`, `best.ckpt.index`, `best.ckpt.meta`, `checkpoint`: TensorFlow / Lingvo checkpoint for `AecSingleInterfering`. - `model.tflite`: Dynamic-range quantized TensorFlow Lite (`.tflite`) FlatBuffer model for on-device inference. - `evaluation_metrics.json`: Verified evaluation results (`test-clean` and `test-other`) for `AecSingleInterfering`. --- ## How to Use ### 1. Install `textual-echo-cancellation` ```bash pip3 install textual-echo-cancellation huggingface_hub ``` ### 2. Download the Model from Hugging Face ```python from huggingface_hub import snapshot_download model_dir = snapshot_download(repo_id="wq2012/aec_single_interfering") print("Downloaded model to:", model_dir) ``` ### 3. Run Inference via CLI (`scripts/inference.py`) ```bash python3 -m scripts.inference \ --model AecSingleInterfering \ --checkpoint_path "${MODEL_DIR}/best.ckpt" \ --mixed_wav /path/to/mixed_input.wav \ --interfering_wav /path/to/reference_tts.wav \ --output_wav /tmp/enhanced_clean.wav ``` ### 4. Run Inference via Python API ```python import os from huggingface_hub import snapshot_download from tec import inference model_dir = snapshot_download(repo_id="wq2012/aec_single_interfering") ckpt_path = os.path.join(model_dir, "best.ckpt") result = inference.run_inference_on_wav( model_name="AecSingleInterfering", mixed_wav_path="/path/to/mixed_input.wav", interfering_text="currently in mountain view it is 72 degrees", checkpoint_path=ckpt_path, output_wav_path="/tmp/enhanced_clean.wav", ) print("Enhanced log-Mel spectrogram shape:", result["predicted_mel"].shape) ``` ### 5. On-Device Inference with Quantized TFLite (`model.tflite`) ```python import os import numpy as np import tensorflow as tf from huggingface_hub import snapshot_download model_dir = snapshot_download(repo_id="wq2012/aec_single_interfering") tflite_path = os.path.join(model_dir, "model.tflite") interpreter = tf.lite.Interpreter(model_path=tflite_path) interpreter.allocate_tensors() for detail in interpreter.get_input_details(): interpreter.set_tensor( detail["index"], np.zeros(detail["shape"], dtype=detail["dtype"]) ) interpreter.invoke() output_details = interpreter.get_output_details() enhanced_mel = interpreter.get_tensor(output_details[0]["index"]) print("TFLite predicted log-Mel shape:", enhanced_mel.shape) ``` --- ## Citation If you use this model or the `textual-echo-cancellation` library in your research, please cite the original paper: ```bibtex @inproceedings{ding2021textual, title={Textual Echo Cancellation}, author={Ding, Shaojin and Jia, Ye and Hu, Ke and Wang, Quan}, booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)}, pages={666--673}, year={2021}, organization={IEEE} } ```