File size: 7,035 Bytes
364eb0d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 | ---
license: apache-2.0
library_name: lingvo
tags:
- audio
- speech-enhancement
- echo-cancellation
- textual-echo-cancellation
- tensorflow
- tflite
datasets:
- libritts
- vctk
metrics:
- wer
- mcd
arxiv: 2008.06006
---
# AEC-Seq2seq Neural Baseline β Multiple Interfering Voices (`AecMultiInterfering`)
Multi-source attention sequence-to-sequence **Acoustic Echo Cancellation (`AEC-Seq2seq`)** neural baseline model trained on 24 kHz **LibriTTS** user speech mixed at 0 dB SNR with reverberant **CSTR VCTK** (109-speaker TTS interference). Uses dual audio encoders (`SpeechEncoderV1`) on both the microphone mixture and the full reference TTS playback audio (~187β230 KB side input).
- **GitHub Repository**: [https://github.com/wq2012/tec](https://github.com/wq2012/tec)
- **PyPI Package**: [`textual-echo-cancellation`](https://pypi.org/project/textual-echo-cancellation/)
- **Paper**: [Textual Echo Cancellation (IEEE SLT 2021, arXiv:2008.06006v4)](https://arxiv.org/pdf/2008.06006)
- **Audio Demo Page**: [https://google.github.io/speaker-id/publications/TEC/](https://google.github.io/speaker-id/publications/TEC/)
> **Open-Source Reproduction Notice**: This pretrained model was trained using the standalone open-source reproduction library ([`wq2012/tec`](https://github.com/wq2012/tec)) built on [`lingvo`](https://github.com/tensorflow/lingvo) and `tensorflow`. It does not use Google's proprietary internal codebase or internal training data infrastructure.
---
## Model Performance
### 1. Open-Source Reproduction Evaluation (`AecMultiInterfering`)
Evaluated on 24 kHz **LibriTTS** (`test-clean`, `test-other`) mixed at 0 dB SNR with synthetic room impulse responses (RT60 = 0.25 s) under the **Multiple interfering voices (LibriTTS + VCTK)** condition, using `Qwen3-ASR-0.6B-F16` via `audio.cpp` for ASR WER scoring and 13-MFCC Dynamic Time Warping for MCD:
| Metric | `test-clean` | `test-other` |Paper Reference (`AEC-Seq2seq`) |
| :--- | :---: | :---: | :---: |
| **WER (%) β** | **8.54%** (17/199) | **22.22%** (54/243) | 6.90% / 19.8% |
| **MCD (dB) β** | **7.80 dB** | **7.88 dB** | 5.04 dB / 5.72 dB |
| **Side Input Size (KB) β** | **186.922 KB** | **206.759 KB** | 230 KB |
| **Computational Complexity β** | **8.62 GFLOPS** | **8.62 GFLOPS** | 8.62 GFLOPS |
### 2. Full Comparison Across All Pretrained Models & Baselines
| Condition | Method | Hugging Face Model | WER (%) test-clean β | WER (%) test-other β | MCD (dB) test-clean β | MCD (dB) test-other β | Side Input test-clean (KB) β |
| :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
| **Single interfering voice** *(LibriTTS + LJSpeech)* | `GroundTruth` | β | 3.52 | 6.78 | 0.00 | 0.00 | 0.000 |
| | `MicrophoneSignal` | β | 90.45 | 114.12 | 12.86 | 14.61 | 0.000 |
| | `NlmsAec` (AEC-NLMS) | β | 88.44 | 107.34 | 12.80 | 14.48 | 243.465 |
| | `NoSideInputSingleInterfering` | [`wq2012/vanilla_seq2seq_single_interfering`](https://huggingface.co/wq2012/vanilla_seq2seq_single_interfering) | 45.23 | 91.53 | 9.58 | 11.34 | 0.000 |
| | `AecSingleInterfering` | [`wq2012/aec_single_interfering`](https://huggingface.co/wq2012/aec_single_interfering) | 12.06 | 23.16 | 8.85 | 9.86 | 243.465 |
| | **`TecSingleInterfering`** | **[`wq2012/tec_single_interfering`](https://huggingface.co/wq2012/tec_single_interfering)** | **21.61** | **46.89** | **8.24** | **9.28** | **0.076** |
| **Multiple interfering voices** *(LibriTTS + VCTK)* | `GroundTruth` | β | 5.03 | 7.82 | 0.00 | 0.00 | 0.000 |
| | `MicrophoneSignal` | β | 34.17 | 48.97 | 7.67 | 7.70 | 0.000 |
| | `NlmsAec` (AEC-NLMS) | β | 28.64 | 34.98 | 7.92 | 8.40 | 186.922 |
| | `NoSideInputMultiInterfering` | [`wq2012/vanilla_seq2seq_multi_interfering`](https://huggingface.co/wq2012/vanilla_seq2seq_multi_interfering) | 31.16 | 42.39 | 7.93 | 8.72 | 0.000 |
| | `AecMultiInterfering` | [`wq2012/aec_multi_interfering`](https://huggingface.co/wq2012/aec_multi_interfering) | 8.54 | 22.22 | 7.80 | 7.88 | 186.922 |
| | **`TecMultiInterfering`** | **[`wq2012/tec_multi_interfering`](https://huggingface.co/wq2012/tec_multi_interfering)** | **26.63** | **45.27** | **7.96** | **8.40** | **0.037** |
---
## Files in This Repository
- `best.ckpt.data-00000-of-00001`, `best.ckpt.index`, `best.ckpt.meta`, `checkpoint`: TensorFlow / Lingvo checkpoint for `AecMultiInterfering`.
- `model.tflite`: Dynamic-range quantized TensorFlow Lite (`.tflite`) FlatBuffer model for on-device inference.
- `evaluation_metrics.json`: Verified evaluation results (`test-clean` and `test-other`) for `AecMultiInterfering`.
---
## How to Use
### 1. Install `textual-echo-cancellation`
```bash
pip3 install textual-echo-cancellation huggingface_hub
```
### 2. Download the Model from Hugging Face
```python
from huggingface_hub import snapshot_download
model_dir = snapshot_download(repo_id="wq2012/aec_multi_interfering")
print("Downloaded model to:", model_dir)
```
### 3. Run Inference via CLI (`scripts/inference.py`)
```bash
python3 -m scripts.inference \
--model AecMultiInterfering \
--checkpoint_path "${MODEL_DIR}/best.ckpt" \
--mixed_wav /path/to/mixed_input.wav \
--interfering_wav /path/to/reference_tts.wav \
--output_wav /tmp/enhanced_clean.wav
```
### 4. Run Inference via Python API
```python
import os
from huggingface_hub import snapshot_download
from tec import inference
model_dir = snapshot_download(repo_id="wq2012/aec_multi_interfering")
ckpt_path = os.path.join(model_dir, "best.ckpt")
result = inference.run_inference_on_wav(
model_name="AecMultiInterfering",
mixed_wav_path="/path/to/mixed_input.wav",
interfering_text="currently in mountain view it is 72 degrees",
checkpoint_path=ckpt_path,
output_wav_path="/tmp/enhanced_clean.wav",
)
print("Enhanced log-Mel spectrogram shape:", result["predicted_mel"].shape)
```
### 5. On-Device Inference with Quantized TFLite (`model.tflite`)
```python
import os
import numpy as np
import tensorflow as tf
from huggingface_hub import snapshot_download
model_dir = snapshot_download(repo_id="wq2012/aec_multi_interfering")
tflite_path = os.path.join(model_dir, "model.tflite")
interpreter = tf.lite.Interpreter(model_path=tflite_path)
interpreter.allocate_tensors()
for detail in interpreter.get_input_details():
interpreter.set_tensor(
detail["index"], np.zeros(detail["shape"], dtype=detail["dtype"])
)
interpreter.invoke()
output_details = interpreter.get_output_details()
enhanced_mel = interpreter.get_tensor(output_details[0]["index"])
print("TFLite predicted log-Mel shape:", enhanced_mel.shape)
```
---
## Citation
If you use this model or the `textual-echo-cancellation` library in your research, please cite the original paper:
```bibtex
@inproceedings{ding2021textual,
title={Textual Echo Cancellation},
author={Ding, Shaojin and Jia, Ye and Hu, Ke and Wang, Quan},
booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)},
pages={666--673},
year={2021},
organization={IEEE}
}
```
|