File size: 10,958 Bytes
c24c082 627141b c24c082 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 | # Training recipe: Matcha-TTS-PL
The complete, flattened procedure that produces the released model: a Polish Matcha-TTS acoustic model
(base training, then fine-tuning on 8 readers) plus a HiFi-GAN vocoder fine-tuned to it. Everything below uses public data and public
code and runs on one consumer or data-centre GPU in about 6 GPU-hours (RTX 4090: base 3.5 h, target 1 h,
vocoder 1.7 h). Intermediate experiments are not part of this document.
## 0. Ingredients
| item | source | licence |
|---|---|---|
| Matcha-TTS code | github.com/shivammehta25/Matcha-TTS | MIT |
| Warm-start checkpoint | Matcha-TTS release `matcha_vctk.ckpt` (trained on VCTK) | MIT weights; VCTK corpus CC BY 4.0 |
| Vocoder | HiFi-GAN universal v1 `g_02500000` + discriminator `do_02500000` (Matcha-TTS release / HF mirror `AlexAlexBabarika/hifigan-universal-v1`) | MIT |
| HiFi-GAN training code | github.com/jik876/hifi-gan | MIT |
| Wolne Lektury audiobooks | wolnelektury.pl (repack `datadriven-company/WolneLektury-TTS-Polish` on Hugging Face) | CC BY-SA 3.0 PL (attribution: author, title, reader, director) |
| AZON spontaneous speech (`pwr-azon_spont`) | Politechnika Wrocławska | CC BY-SA 4.0 |
| Phonemizer | espeak-ng 1.52 via `phonemizer` | GPL-3.0 (runtime dependency only) |
| Speech recogniser for data filtering | faster-whisper `large-v3-turbo` | MIT |
Tools referenced below live in the `scripts/` directory of the playground repository
([machinekind/tts-pl-playground](https://github.com/machinekind/tts-pl-playground)); `matcha_patch/` holds the Polish
cleaner and configs. Apply `python scripts/patch_matcha.py <repo>` inside the Matcha-TTS clone once (idempotent).
## 1. Environment
```bash
git clone https://github.com/shivammehta25/Matcha-TTS.git
python3.12 -m venv .venv && . .venv/bin/activate
pip install torch torchaudio # CUDA build for training, CPU build is enough for inference
pip install Cython numpy && (cd Matcha-TTS && pip install -e . --no-deps --no-build-isolation)
pip install -r matcha_patch/requirements_min.txt librosa soundfile phonemizer faster-whisper huggingface_hub tensorboard
(cd Matcha-TTS && python ../scripts/patch_matcha.py ..)
export PHONEMIZER_ESPEAK_LIBRARY=/path/to/libespeak-ng.so # macOS: /opt/homebrew/lib/libespeak-ng.dylib
```
What the patch changes in Matcha-TTS:
- `polish_cleaners`: espeak-ng `pl` phonemization with punctuation preserved, text normalisation (numbers, abbreviations).
- Symbol table: adds U+0303 (combining tilde, nasal vowels) → `n_vocab = 179`.
- **Style token**: optional 4th filelist column `path|speaker|text|style`; `MatchaTTS(n_styles=K)` adds a zero-initialised
embedding table `E_k ∈ R^{K×64}` summed with the speaker embedding before the encoder and the decoder.
- Monotonic alignment search through pinned memory; `torchaudio.load` → `soundfile`; DataLoader `spawn`; numpy 2 fixes.
## 2. Data
### 2.1 Ingest
- `scripts/ingest_wolnelektury.py`: download audiobooks, segment on silences to 1–15 s, align text, write
`wavs/*.wav` (22.05 kHz mono, peak −0.45 dBFS), `piper_metadata.csv` (`file|reader|text`) and `sources.json`
(book → author, title, reader, director, licence URL) for attribution.
- `scripts/ingest_azon.py`: unpack, resample to 22.05 kHz, keep speaker ids.
### 2.2 Per-clip statistics and text/audio agreement
```bash
python scripts/clip_stats.py data/wl --out data_v2/clip_stats_wl.csv --workers 10
python scripts/clip_stats.py data/azon --out data_v2/clip_stats_azon.csv
python scripts/wl_book_meta.py data/wl/wl_api_cache.json --out data_v2/wl_books.json # genre per book (WL API)
python scripts/whisper_check.py data/wl --out data_v2/whisper_wl.csv --model large-v3-turbo --device cuda
```
Per clip: duration, RMS, silence fraction, F0 median and spread (pyin), UTMOS (`tarepan/SpeechMOS`), DNSMOS,
characters per second, and the character error rate between the label and a Whisper transcript. Clips with CER > 0.2
are text/audio mismatches (mis-segmented or mislabelled) and are dropped: about a quarter of the speaking-rate
outliers turned out to be mislabelled this way.
### 2.3 Build the training sets
```bash
python scripts/build_mix_v3.py --out data_v2 --q-oversample 4
```
Filters: 1–15 s, UTMOS ≥ 3.0 (AZON ≥ 2.8), DNSMOS ≥ 3.3, Whisper CER ≤ 0.2, Wolne Lektury **prose only** (verse, drama
and fables excluded by genre). Reader selection is by **consistency**, not hours: for each reader
`score = z(UTMOS std) + z(DNSMOS std) + z(spread of per-clip F0 median) − z(UTMOS mean)`, lower is better; the 15 most
consistent readers form the base set, the 8 best are the fine-tune targets (ids 0–7). Speaker-balanced sampling
∝ √hours (1–4×), per-speaker validation split of 2 %.
**Question labels.** Audiobook readers read most questions with a falling contour, so a raw "?" would teach "slightly
less fall". For clips ending in "?", the final-0.9 s pitch slope decides: rising → keep "?" and oversample 4×;
falling but starting with a wh-word (co, kto, gdzie, kiedy, dlaczego, jak, ile, który…) → keep "?" (falling is correct
Polish); falling otherwise → relabel "?" as "." so the label matches the audio.
**Style tokens.** 18 designed tokens = pitch range {flat, mid, wide} × speaking rate {slow, normal, fast} × question
{no, yes}, assigned per clip from the reader-relative F0 spread, rate and the question flag (`style_map.json`; the
neutral token is the one the map names `neutral`). Written as the 4th filelist column.
Outputs: `data_v2/mix_v3` (base: 15 WL readers + 5 AZON speakers, 9.5k clips, ≈ 27 h) and `data_v2/mix_v3_target`
(8 readers, 5.2k clips, ≈ 14 h). The filelists are reproducible from the public sources with the commands above.
## 3. Acoustic model (Matcha-TTS)
Common settings: batch 64, bf16, Adam, `out_size` per Matcha default, checkpoints every N steps plus `final.ckpt`,
validation every 5 epochs, mel statistics computed per dataset (`matcha.utils.generate_data_statistics`).
0.25 s/step on an H100, 0.31 s/step on an RTX 4090.
| stage | data | init | steps | lr |
|---|---|---|---|---|
| base | `mix_v3` (20 speakers, 18 style tokens) | `matcha_vctk.ckpt`, weights only; speaker, style and symbol tables re-initialised | 40 000 | 1e-4 |
| target | `mix_v3_target` (8 readers) | base `final.ckpt` | 12 000 | 5e-5 |
```bash
MIX=mix_v3 RUN_NAME=matcha_v3_base STYLE=1 MAX_STEPS=40000 CKPT_EVERY_STEPS=4000 \
WARM_CKPT=checkpoints/matcha_vctk.ckpt scripts/train_matcha.sh
MIX=mix_v3_target RUN_NAME=matcha_v3_target STYLE=1 MAX_STEPS=12000 CKPT_EVERY_STEPS=2000 \
WARM_CKPT=runs/matcha_v3_base/checkpoints/final.ckpt EXTRA="model.optimizer.lr=5e-5" scripts/train_matcha.sh
```
Warm start is weights-only with a shape-aware partial copy (`scripts/train_matcha_warm.py`): tables that changed size
(speaker, style, symbol embeddings) are copied row by row where shapes overlap, the rest is re-initialised.
Hydra rejects `=` inside override values, so checkpoint files named `step_step=N.ckpt` must be copied to a plain name
before being passed as `+warm_ckpt=`.
The base checkpoint (15 readers + AZON, 40k steps) is not published; to add a new voice, fine-tune the target checkpoint on one to two hours of clean recordings of one speaker with the
target command above, pointing `MIX` at your own filelist.
## 4. Vocoder (HiFi-GAN fine-tuned to the acoustic model)
Matcha's predicted mels are smoother than real ones (lower harmonic contrast, ≈ 7 dB less energy above 5 kHz). A
vocoder trained only on real mels renders that blur as a phasey second layer and electronic "breaths" in pauses.
Fine-tuning the universal HiFi-GAN on the target model's own mels removes it (the standard "fine-tuning" setting of the
HiFi-GAN paper).
```bash
# teacher-forced mels for every target clip: MAS alignment on the real mel -> decoder (10 ODE steps, T 0.5), exact GT length
python scripts/gen_mels_tf.py --ckpt runs/matcha_v3_target/checkpoints/final.ckpt \
--filelist data_v2/mix_v3_target/matcha_train.txt --out vocoder_ft --max-clips 6000 --val 60 --device cuda
# fine-tune generator + discriminators from g_02500000 / do_02500000
FT_LR=2e-5 FT_DISC_LR_SCALE=0.5 FT_WARMUP=2000 FT_BATCH=16 FT_CKPT_EVERY=5000 \
scripts/finetune_hifigan.sh vocoder_ft runs/hifigan_ft 93 # 93 epochs × 323 steps ≈ 30k steps, 0.2 s/step
```
`scripts/patch_hifigan_train.py` adapts the official trainer: it restores the learning rate after loading the optimizer
state from `do_02500000` (which otherwise silently overrides it), trains the generator alone on the mel loss for the
first 2 000 steps before enabling the adversarial and feature-matching losses, and fixes torch ≥ 2 / librosa ≥ 0.10
incompatibilities. Generator lr 2e-5, discriminators 1e-5, batch 16, segment 8192, 30k steps. The released vocoder is
the 30k-step generator; checkpoints from 15k on sound alike.
## 5. Evaluation
`scripts/synth_samples.py` (10 conversational sentences, 4–10 ODE steps, temperature 0.5, length scale 0.9) →
`scripts/eval_synth.py`: Whisper large-v3 WER/CER, UTMOS, F0 spread in semitones, speaking rate, silence ratio.
Note that UTMOS does not register the vocoder artefacts described in §4 (it even scores the fine-tuned vocoder slightly
lower); the vocoder decision was made by listening, with `scripts/diag_vocoder_copy.py` (copy-synthesis of real clips)
and `scripts/diag_phone_artifacts.py` (per-phoneme roughness) as diagnostics.
## 6. Inference
- Recommended settings: 4 ODE steps, temperature 0.5, length scale 0.9–1.0, the fine-tuned vocoder, a 10 kHz low-pass
on the output (removes a HiFi-GAN upsampling tone at the Nyquist frequency; `effects.py` in the playground).
- Voice blending: average speaker-embedding rows with weights (`--voice "0.5*6+0.5*3"`). The ONNX exports carry extra
rows with blends already baked in (`scripts/bake_voice.py`), so a blend is just a speaker id.
- Style tokens: an extra embedding added to the speaker embedding; ids and meanings in `style_map.json`.
- ONNX (acoustic model + vocoder in one graph, inputs `x`, `x_lengths`, `scales=[temperature, length_scale]`, `spks`):
`scripts/export_matcha_onnx.py <ckpt> out.onnx --n-timesteps 4 --vocoder-name hifigan_univ_v1 --vocoder-checkpoint-path <fine-tuned g_*>`.
- Measured on an NVIDIA GB10 (PyTorch bf16 + `torch.compile`, batch 1, 4 steps): ≈ 23 ms to first audio, RTF ≈ 0.006.
Apple M-series CPU, ONNX Runtime: ≈ 0.4 s for a 4 s sentence.
## 7. Licences of the result
Weights trained on CC BY-SA material are released under **CC BY-SA 4.0** (compatible with CC BY-SA 3.0 PL).
`ATTRIBUTION.md` lists every Wolne Lektury book (author, title, reader, director, URL) and the AZON corpus; the warm
start (VCTK) and HiFi-GAN are MIT. Readers' voices are personal attributes not covered by the Creative Commons licence;
the release recommends blended voices under neutral names and disclosure of synthetic speech to listeners.
|