Matcha-TTS-PL / RECIPE.md
mcPear's picture
RECIPE: new voice needs one to two hours of clean recordings
627141b verified
|
Raw History Blame Contribute Delete
11 kB
# Training recipe: Matcha-TTS-PL
The complete, flattened procedure that produces the released model: a Polish Matcha-TTS acoustic model
(base training, then fine-tuning on 8 readers) plus a HiFi-GAN vocoder fine-tuned to it. Everything below uses public data and public
code and runs on one consumer or data-centre GPU in about 6 GPU-hours (RTX 4090: base 3.5 h, target 1 h,
vocoder 1.7 h). Intermediate experiments are not part of this document.
## 0. Ingredients
| item | source | licence |
|---|---|---|
| Matcha-TTS code | github.com/shivammehta25/Matcha-TTS | MIT |
| Warm-start checkpoint | Matcha-TTS release `matcha_vctk.ckpt` (trained on VCTK) | MIT weights; VCTK corpus CC BY 4.0 |
| Vocoder | HiFi-GAN universal v1 `g_02500000` + discriminator `do_02500000` (Matcha-TTS release / HF mirror `AlexAlexBabarika/hifigan-universal-v1`) | MIT |
| HiFi-GAN training code | github.com/jik876/hifi-gan | MIT |
| Wolne Lektury audiobooks | wolnelektury.pl (repack `datadriven-company/WolneLektury-TTS-Polish` on Hugging Face) | CC BY-SA 3.0 PL (attribution: author, title, reader, director) |
| AZON spontaneous speech (`pwr-azon_spont`) | Politechnika WrocΕ‚awska | CC BY-SA 4.0 |
| Phonemizer | espeak-ng 1.52 via `phonemizer` | GPL-3.0 (runtime dependency only) |
| Speech recogniser for data filtering | faster-whisper `large-v3-turbo` | MIT |
Tools referenced below live in the `scripts/` directory of the playground repository
([machinekind/tts-pl-playground](https://github.com/machinekind/tts-pl-playground)); `matcha_patch/` holds the Polish
cleaner and configs. Apply `python scripts/patch_matcha.py <repo>` inside the Matcha-TTS clone once (idempotent).
## 1. Environment
```bash
git clone https://github.com/shivammehta25/Matcha-TTS.git
python3.12 -m venv .venv && . .venv/bin/activate
pip install torch torchaudio # CUDA build for training, CPU build is enough for inference
pip install Cython numpy && (cd Matcha-TTS && pip install -e . --no-deps --no-build-isolation)
pip install -r matcha_patch/requirements_min.txt librosa soundfile phonemizer faster-whisper huggingface_hub tensorboard
(cd Matcha-TTS && python ../scripts/patch_matcha.py ..)
export PHONEMIZER_ESPEAK_LIBRARY=/path/to/libespeak-ng.so # macOS: /opt/homebrew/lib/libespeak-ng.dylib
```
What the patch changes in Matcha-TTS:
- `polish_cleaners`: espeak-ng `pl` phonemization with punctuation preserved, text normalisation (numbers, abbreviations).
- Symbol table: adds U+0303 (combining tilde, nasal vowels) β†’ `n_vocab = 179`.
- **Style token**: optional 4th filelist column `path|speaker|text|style`; `MatchaTTS(n_styles=K)` adds a zero-initialised
embedding table `E_k ∈ R^{KΓ—64}` summed with the speaker embedding before the encoder and the decoder.
- Monotonic alignment search through pinned memory; `torchaudio.load` β†’ `soundfile`; DataLoader `spawn`; numpy 2 fixes.
## 2. Data
### 2.1 Ingest
- `scripts/ingest_wolnelektury.py`: download audiobooks, segment on silences to 1–15 s, align text, write
`wavs/*.wav` (22.05 kHz mono, peak βˆ’0.45 dBFS), `piper_metadata.csv` (`file|reader|text`) and `sources.json`
(book β†’ author, title, reader, director, licence URL) for attribution.
- `scripts/ingest_azon.py`: unpack, resample to 22.05 kHz, keep speaker ids.
### 2.2 Per-clip statistics and text/audio agreement
```bash
python scripts/clip_stats.py data/wl --out data_v2/clip_stats_wl.csv --workers 10
python scripts/clip_stats.py data/azon --out data_v2/clip_stats_azon.csv
python scripts/wl_book_meta.py data/wl/wl_api_cache.json --out data_v2/wl_books.json # genre per book (WL API)
python scripts/whisper_check.py data/wl --out data_v2/whisper_wl.csv --model large-v3-turbo --device cuda
```
Per clip: duration, RMS, silence fraction, F0 median and spread (pyin), UTMOS (`tarepan/SpeechMOS`), DNSMOS,
characters per second, and the character error rate between the label and a Whisper transcript. Clips with CER > 0.2
are text/audio mismatches (mis-segmented or mislabelled) and are dropped: about a quarter of the speaking-rate
outliers turned out to be mislabelled this way.
### 2.3 Build the training sets
```bash
python scripts/build_mix_v3.py --out data_v2 --q-oversample 4
```
Filters: 1–15 s, UTMOS β‰₯ 3.0 (AZON β‰₯ 2.8), DNSMOS β‰₯ 3.3, Whisper CER ≀ 0.2, Wolne Lektury **prose only** (verse, drama
and fables excluded by genre). Reader selection is by **consistency**, not hours: for each reader
`score = z(UTMOS std) + z(DNSMOS std) + z(spread of per-clip F0 median) βˆ’ z(UTMOS mean)`, lower is better; the 15 most
consistent readers form the base set, the 8 best are the fine-tune targets (ids 0–7). Speaker-balanced sampling
∝ √hours (1–4Γ—), per-speaker validation split of 2 %.
**Question labels.** Audiobook readers read most questions with a falling contour, so a raw "?" would teach "slightly
less fall". For clips ending in "?", the final-0.9 s pitch slope decides: rising β†’ keep "?" and oversample 4Γ—;
falling but starting with a wh-word (co, kto, gdzie, kiedy, dlaczego, jak, ile, ktΓ³ry…) β†’ keep "?" (falling is correct
Polish); falling otherwise β†’ relabel "?" as "." so the label matches the audio.
**Style tokens.** 18 designed tokens = pitch range {flat, mid, wide} Γ— speaking rate {slow, normal, fast} Γ— question
{no, yes}, assigned per clip from the reader-relative F0 spread, rate and the question flag (`style_map.json`; the
neutral token is the one the map names `neutral`). Written as the 4th filelist column.
Outputs: `data_v2/mix_v3` (base: 15 WL readers + 5 AZON speakers, 9.5k clips, β‰ˆ 27 h) and `data_v2/mix_v3_target`
(8 readers, 5.2k clips, β‰ˆ 14 h). The filelists are reproducible from the public sources with the commands above.
## 3. Acoustic model (Matcha-TTS)
Common settings: batch 64, bf16, Adam, `out_size` per Matcha default, checkpoints every N steps plus `final.ckpt`,
validation every 5 epochs, mel statistics computed per dataset (`matcha.utils.generate_data_statistics`).
0.25 s/step on an H100, 0.31 s/step on an RTX 4090.
| stage | data | init | steps | lr |
|---|---|---|---|---|
| base | `mix_v3` (20 speakers, 18 style tokens) | `matcha_vctk.ckpt`, weights only; speaker, style and symbol tables re-initialised | 40 000 | 1e-4 |
| target | `mix_v3_target` (8 readers) | base `final.ckpt` | 12 000 | 5e-5 |
```bash
MIX=mix_v3 RUN_NAME=matcha_v3_base STYLE=1 MAX_STEPS=40000 CKPT_EVERY_STEPS=4000 \
WARM_CKPT=checkpoints/matcha_vctk.ckpt scripts/train_matcha.sh
MIX=mix_v3_target RUN_NAME=matcha_v3_target STYLE=1 MAX_STEPS=12000 CKPT_EVERY_STEPS=2000 \
WARM_CKPT=runs/matcha_v3_base/checkpoints/final.ckpt EXTRA="model.optimizer.lr=5e-5" scripts/train_matcha.sh
```
Warm start is weights-only with a shape-aware partial copy (`scripts/train_matcha_warm.py`): tables that changed size
(speaker, style, symbol embeddings) are copied row by row where shapes overlap, the rest is re-initialised.
Hydra rejects `=` inside override values, so checkpoint files named `step_step=N.ckpt` must be copied to a plain name
before being passed as `+warm_ckpt=`.
The base checkpoint (15 readers + AZON, 40k steps) is not published; to add a new voice, fine-tune the target checkpoint on one to two hours of clean recordings of one speaker with the
target command above, pointing `MIX` at your own filelist.
## 4. Vocoder (HiFi-GAN fine-tuned to the acoustic model)
Matcha's predicted mels are smoother than real ones (lower harmonic contrast, β‰ˆ 7 dB less energy above 5 kHz). A
vocoder trained only on real mels renders that blur as a phasey second layer and electronic "breaths" in pauses.
Fine-tuning the universal HiFi-GAN on the target model's own mels removes it (the standard "fine-tuning" setting of the
HiFi-GAN paper).
```bash
# teacher-forced mels for every target clip: MAS alignment on the real mel -> decoder (10 ODE steps, T 0.5), exact GT length
python scripts/gen_mels_tf.py --ckpt runs/matcha_v3_target/checkpoints/final.ckpt \
--filelist data_v2/mix_v3_target/matcha_train.txt --out vocoder_ft --max-clips 6000 --val 60 --device cuda
# fine-tune generator + discriminators from g_02500000 / do_02500000
FT_LR=2e-5 FT_DISC_LR_SCALE=0.5 FT_WARMUP=2000 FT_BATCH=16 FT_CKPT_EVERY=5000 \
scripts/finetune_hifigan.sh vocoder_ft runs/hifigan_ft 93 # 93 epochs Γ— 323 steps β‰ˆ 30k steps, 0.2 s/step
```
`scripts/patch_hifigan_train.py` adapts the official trainer: it restores the learning rate after loading the optimizer
state from `do_02500000` (which otherwise silently overrides it), trains the generator alone on the mel loss for the
first 2 000 steps before enabling the adversarial and feature-matching losses, and fixes torch β‰₯ 2 / librosa β‰₯ 0.10
incompatibilities. Generator lr 2e-5, discriminators 1e-5, batch 16, segment 8192, 30k steps. The released vocoder is
the 30k-step generator; checkpoints from 15k on sound alike.
## 5. Evaluation
`scripts/synth_samples.py` (10 conversational sentences, 4–10 ODE steps, temperature 0.5, length scale 0.9) β†’
`scripts/eval_synth.py`: Whisper large-v3 WER/CER, UTMOS, F0 spread in semitones, speaking rate, silence ratio.
Note that UTMOS does not register the vocoder artefacts described in Β§4 (it even scores the fine-tuned vocoder slightly
lower); the vocoder decision was made by listening, with `scripts/diag_vocoder_copy.py` (copy-synthesis of real clips)
and `scripts/diag_phone_artifacts.py` (per-phoneme roughness) as diagnostics.
## 6. Inference
- Recommended settings: 4 ODE steps, temperature 0.5, length scale 0.9–1.0, the fine-tuned vocoder, a 10 kHz low-pass
on the output (removes a HiFi-GAN upsampling tone at the Nyquist frequency; `effects.py` in the playground).
- Voice blending: average speaker-embedding rows with weights (`--voice "0.5*6+0.5*3"`). The ONNX exports carry extra
rows with blends already baked in (`scripts/bake_voice.py`), so a blend is just a speaker id.
- Style tokens: an extra embedding added to the speaker embedding; ids and meanings in `style_map.json`.
- ONNX (acoustic model + vocoder in one graph, inputs `x`, `x_lengths`, `scales=[temperature, length_scale]`, `spks`):
`scripts/export_matcha_onnx.py <ckpt> out.onnx --n-timesteps 4 --vocoder-name hifigan_univ_v1 --vocoder-checkpoint-path <fine-tuned g_*>`.
- Measured on an NVIDIA GB10 (PyTorch bf16 + `torch.compile`, batch 1, 4 steps): β‰ˆ 23 ms to first audio, RTF β‰ˆ 0.006.
Apple M-series CPU, ONNX Runtime: β‰ˆ 0.4 s for a 4 s sentence.
## 7. Licences of the result
Weights trained on CC BY-SA material are released under **CC BY-SA 4.0** (compatible with CC BY-SA 3.0 PL).
`ATTRIBUTION.md` lists every Wolne Lektury book (author, title, reader, director, URL) and the AZON corpus; the warm
start (VCTK) and HiFi-GAN are MIT. Readers' voices are personal attributes not covered by the Creative Commons licence;
the release recommends blended voices under neutral names and disclosure of synthetic speech to listeners.