wq2012's picture
Upload pretrained tec_single_interfering checkpoint, TFLite model, and Model Card
354accc verified
|
Raw History Blame Contribute Delete
7.1 kB
---
license: apache-2.0
library_name: lingvo
tags:
- audio
- speech-enhancement
- echo-cancellation
- textual-echo-cancellation
- tensorflow
- tflite
datasets:
- libritts
- ljspeech
metrics:
- wer
- mcd
arxiv: 2008.06006
---
# Textual Echo Cancellation (TEC) β€” Single Interfering Voice (`TecSingleInterfering`)
Multi-source attention sequence-to-sequence **Textual Echo Cancellation (TEC)** model trained on 24 kHz **LibriTTS** user speech mixed at 0 dB SNR with reverberant **LJ Speech** (single-speaker TTS interference). Takes the noisy/reverberant microphone spectrogram (`SpeechEncoderV1`) and the lightweight TTS source transcript (`TtsEncoderV2`, < 0.1 KB side input) to reconstruct the clean user speech spectrogram.
- **GitHub Repository**: [https://github.com/wq2012/tec](https://github.com/wq2012/tec)
- **PyPI Package**: [`textual-echo-cancellation`](https://pypi.org/project/textual-echo-cancellation/)
- **Paper**: [Textual Echo Cancellation (IEEE SLT 2021, arXiv:2008.06006v4)](https://arxiv.org/pdf/2008.06006)
- **Audio Demo Page**: [https://google.github.io/speaker-id/publications/TEC/](https://google.github.io/speaker-id/publications/TEC/)
> **Open-Source Reproduction Notice**: This pretrained model was trained using the standalone open-source reproduction library ([`wq2012/tec`](https://github.com/wq2012/tec)) built on [`lingvo`](https://github.com/tensorflow/lingvo) and `tensorflow`. It does not use Google's proprietary internal codebase or internal training data infrastructure.
---
## Model Performance
### 1. Open-Source Reproduction Evaluation (`TecSingleInterfering`)
Evaluated on 24 kHz **LibriTTS** (`test-clean`, `test-other`) mixed at 0 dB SNR with synthetic room impulse responses (RT60 = 0.25 s) under the **Single interfering voice (LibriTTS + LJ Speech)** condition, using `Qwen3-ASR-0.6B-F16` via `audio.cpp` for ASR WER scoring and 13-MFCC Dynamic Time Warping for MCD:
| Metric | `test-clean` | `test-other` |Paper Reference (`TEC (proposed)`) |
| :--- | :---: | :---: | :---: |
| **WER (%) ↓** | **21.61%** (43/199) | **46.89%** (83/177) | 15.5% / 39.8% |
| **MCD (dB) ↓** | **8.24 dB** | **9.28 dB** | 7.51 dB / 8.54 dB |
| **Side Input Size (KB) ↓** | **0.076 KB** | **0.068 KB** | 0.10 KB |
| **Computational Complexity ↓** | **7.27 GFLOPS** | **7.27 GFLOPS** | 7.27 GFLOPS |
### 2. Full Comparison Across All Pretrained Models & Baselines
| Condition | Method | Hugging Face Model | WER (%) test-clean ↓ | WER (%) test-other ↓ | MCD (dB) test-clean ↓ | MCD (dB) test-other ↓ | Side Input test-clean (KB) ↓ |
| :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
| **Single interfering voice** *(LibriTTS + LJSpeech)* | `GroundTruth` | β€” | 3.52 | 6.78 | 0.00 | 0.00 | 0.000 |
| | `MicrophoneSignal` | β€” | 90.45 | 114.12 | 12.86 | 14.61 | 0.000 |
| | `NlmsAec` (AEC-NLMS) | β€” | 88.44 | 107.34 | 12.80 | 14.48 | 243.465 |
| | `NoSideInputSingleInterfering` | [`wq2012/vanilla_seq2seq_single_interfering`](https://huggingface.co/wq2012/vanilla_seq2seq_single_interfering) | 45.23 | 91.53 | 9.58 | 11.34 | 0.000 |
| | `AecSingleInterfering` | [`wq2012/aec_single_interfering`](https://huggingface.co/wq2012/aec_single_interfering) | 12.06 | 23.16 | 8.85 | 9.86 | 243.465 |
| | **`TecSingleInterfering`** | **[`wq2012/tec_single_interfering`](https://huggingface.co/wq2012/tec_single_interfering)** | **21.61** | **46.89** | **8.24** | **9.28** | **0.076** |
| **Multiple interfering voices** *(LibriTTS + VCTK)* | `GroundTruth` | β€” | 5.03 | 7.82 | 0.00 | 0.00 | 0.000 |
| | `MicrophoneSignal` | β€” | 34.17 | 48.97 | 7.67 | 7.70 | 0.000 |
| | `NlmsAec` (AEC-NLMS) | β€” | 28.64 | 34.98 | 7.92 | 8.40 | 186.922 |
| | `NoSideInputMultiInterfering` | [`wq2012/vanilla_seq2seq_multi_interfering`](https://huggingface.co/wq2012/vanilla_seq2seq_multi_interfering) | 31.16 | 42.39 | 7.93 | 8.72 | 0.000 |
| | `AecMultiInterfering` | [`wq2012/aec_multi_interfering`](https://huggingface.co/wq2012/aec_multi_interfering) | 8.54 | 22.22 | 7.80 | 7.88 | 186.922 |
| | **`TecMultiInterfering`** | **[`wq2012/tec_multi_interfering`](https://huggingface.co/wq2012/tec_multi_interfering)** | **26.63** | **45.27** | **7.96** | **8.40** | **0.037** |
---
## Files in This Repository
- `best.ckpt.data-00000-of-00001`, `best.ckpt.index`, `best.ckpt.meta`, `checkpoint`: TensorFlow / Lingvo checkpoint for `TecSingleInterfering`.
- `model.tflite`: Dynamic-range quantized TensorFlow Lite (`.tflite`) FlatBuffer model for on-device inference.
- `evaluation_metrics.json`: Verified evaluation results (`test-clean` and `test-other`) for `TecSingleInterfering`.
---
## How to Use
### 1. Install `textual-echo-cancellation`
```bash
pip3 install textual-echo-cancellation huggingface_hub
```
### 2. Download the Model from Hugging Face
```python
from huggingface_hub import snapshot_download
model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
print("Downloaded model to:", model_dir)
```
### 3. Run Inference via CLI (`scripts/inference.py`)
```bash
python3 -m scripts.inference \
--model TecSingleInterfering \
--checkpoint_path "${MODEL_DIR}/best.ckpt" \
--mixed_wav /path/to/mixed_input.wav \
--interfering_text "currently in mountain view it is 72 degrees" \
--output_wav /tmp/enhanced_clean.wav
```
### 4. Run Inference via Python API
```python
import os
from huggingface_hub import snapshot_download
from tec import inference
model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
ckpt_path = os.path.join(model_dir, "best.ckpt")
result = inference.run_inference_on_wav(
model_name="TecSingleInterfering",
mixed_wav_path="/path/to/mixed_input.wav",
interfering_text="currently in mountain view it is 72 degrees",
checkpoint_path=ckpt_path,
output_wav_path="/tmp/enhanced_clean.wav",
)
print("Enhanced log-Mel spectrogram shape:", result["predicted_mel"].shape)
```
### 5. On-Device Inference with Quantized TFLite (`model.tflite`)
```python
import os
import numpy as np
import tensorflow as tf
from huggingface_hub import snapshot_download
model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
tflite_path = os.path.join(model_dir, "model.tflite")
interpreter = tf.lite.Interpreter(model_path=tflite_path)
interpreter.allocate_tensors()
for detail in interpreter.get_input_details():
interpreter.set_tensor(
detail["index"], np.zeros(detail["shape"], dtype=detail["dtype"])
)
interpreter.invoke()
output_details = interpreter.get_output_details()
enhanced_mel = interpreter.get_tensor(output_details[0]["index"])
print("TFLite predicted log-Mel shape:", enhanced_mel.shape)
```
---
## Citation
If you use this model or the `textual-echo-cancellation` library in your research, please cite the original paper:
```bibtex
@inproceedings{ding2021textual,
title={Textual Echo Cancellation},
author={Ding, Shaojin and Jia, Ye and Hu, Ke and Wang, Quan},
booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)},
pages={666--673},
year={2021},
organization={IEEE}
}
```