|
Download RECIPE.md from machinekind/Matcha-TTS-PL: direct link, hf CLI and curl.
- Browser
- Download file 11 kB
-
https://huggingface.co/machinekind/Matcha-TTS-PL/resolve/main/RECIPE.md
- Command line
-
hf download hf://machinekind/Matcha-TTS-PL/RECIPE.md
-
curl -L -o RECIPE.md https://huggingface.co/machinekind/Matcha-TTS-PL/resolve/main/RECIPE.md
11 kB
| # Training recipe: Matcha-TTS-PL | |
| The complete, flattened procedure that produces the released model: a Polish Matcha-TTS acoustic model | |
| (base training, then fine-tuning on 8 readers) plus a HiFi-GAN vocoder fine-tuned to it. Everything below uses public data and public | |
| code and runs on one consumer or data-centre GPU in about 6 GPU-hours (RTX 4090: base 3.5 h, target 1 h, | |
| vocoder 1.7 h). Intermediate experiments are not part of this document. | |
| ## 0. Ingredients | |
| | item | source | licence | | |
| |---|---|---| | |
| | Matcha-TTS code | github.com/shivammehta25/Matcha-TTS | MIT | | |
| | Warm-start checkpoint | Matcha-TTS release `matcha_vctk.ckpt` (trained on VCTK) | MIT weights; VCTK corpus CC BY 4.0 | | |
| | Vocoder | HiFi-GAN universal v1 `g_02500000` + discriminator `do_02500000` (Matcha-TTS release / HF mirror `AlexAlexBabarika/hifigan-universal-v1`) | MIT | | |
| | HiFi-GAN training code | github.com/jik876/hifi-gan | MIT | | |
| | Wolne Lektury audiobooks | wolnelektury.pl (repack `datadriven-company/WolneLektury-TTS-Polish` on Hugging Face) | CC BY-SA 3.0 PL (attribution: author, title, reader, director) | | |
| | AZON spontaneous speech (`pwr-azon_spont`) | Politechnika WrocΕawska | CC BY-SA 4.0 | | |
| | Phonemizer | espeak-ng 1.52 via `phonemizer` | GPL-3.0 (runtime dependency only) | | |
| | Speech recogniser for data filtering | faster-whisper `large-v3-turbo` | MIT | | |
| Tools referenced below live in the `scripts/` directory of the playground repository | |
| ([machinekind/tts-pl-playground](https://github.com/machinekind/tts-pl-playground)); `matcha_patch/` holds the Polish | |
| cleaner and configs. Apply `python scripts/patch_matcha.py <repo>` inside the Matcha-TTS clone once (idempotent). | |
| ## 1. Environment | |
| ```bash | |
| git clone https://github.com/shivammehta25/Matcha-TTS.git | |
| python3.12 -m venv .venv && . .venv/bin/activate | |
| pip install torch torchaudio # CUDA build for training, CPU build is enough for inference | |
| pip install Cython numpy && (cd Matcha-TTS && pip install -e . --no-deps --no-build-isolation) | |
| pip install -r matcha_patch/requirements_min.txt librosa soundfile phonemizer faster-whisper huggingface_hub tensorboard | |
| (cd Matcha-TTS && python ../scripts/patch_matcha.py ..) | |
| export PHONEMIZER_ESPEAK_LIBRARY=/path/to/libespeak-ng.so # macOS: /opt/homebrew/lib/libespeak-ng.dylib | |
| ``` | |
| What the patch changes in Matcha-TTS: | |
| - `polish_cleaners`: espeak-ng `pl` phonemization with punctuation preserved, text normalisation (numbers, abbreviations). | |
| - Symbol table: adds U+0303 (combining tilde, nasal vowels) β `n_vocab = 179`. | |
| - **Style token**: optional 4th filelist column `path|speaker|text|style`; `MatchaTTS(n_styles=K)` adds a zero-initialised | |
| embedding table `E_k β R^{KΓ64}` summed with the speaker embedding before the encoder and the decoder. | |
| - Monotonic alignment search through pinned memory; `torchaudio.load` β `soundfile`; DataLoader `spawn`; numpy 2 fixes. | |
| ## 2. Data | |
| ### 2.1 Ingest | |
| - `scripts/ingest_wolnelektury.py`: download audiobooks, segment on silences to 1β15 s, align text, write | |
| `wavs/*.wav` (22.05 kHz mono, peak β0.45 dBFS), `piper_metadata.csv` (`file|reader|text`) and `sources.json` | |
| (book β author, title, reader, director, licence URL) for attribution. | |
| - `scripts/ingest_azon.py`: unpack, resample to 22.05 kHz, keep speaker ids. | |
| ### 2.2 Per-clip statistics and text/audio agreement | |
| ```bash | |
| python scripts/clip_stats.py data/wl --out data_v2/clip_stats_wl.csv --workers 10 | |
| python scripts/clip_stats.py data/azon --out data_v2/clip_stats_azon.csv | |
| python scripts/wl_book_meta.py data/wl/wl_api_cache.json --out data_v2/wl_books.json # genre per book (WL API) | |
| python scripts/whisper_check.py data/wl --out data_v2/whisper_wl.csv --model large-v3-turbo --device cuda | |
| ``` | |
| Per clip: duration, RMS, silence fraction, F0 median and spread (pyin), UTMOS (`tarepan/SpeechMOS`), DNSMOS, | |
| characters per second, and the character error rate between the label and a Whisper transcript. Clips with CER > 0.2 | |
| are text/audio mismatches (mis-segmented or mislabelled) and are dropped: about a quarter of the speaking-rate | |
| outliers turned out to be mislabelled this way. | |
| ### 2.3 Build the training sets | |
| ```bash | |
| python scripts/build_mix_v3.py --out data_v2 --q-oversample 4 | |
| ``` | |
| Filters: 1β15 s, UTMOS β₯ 3.0 (AZON β₯ 2.8), DNSMOS β₯ 3.3, Whisper CER β€ 0.2, Wolne Lektury **prose only** (verse, drama | |
| and fables excluded by genre). Reader selection is by **consistency**, not hours: for each reader | |
| `score = z(UTMOS std) + z(DNSMOS std) + z(spread of per-clip F0 median) β z(UTMOS mean)`, lower is better; the 15 most | |
| consistent readers form the base set, the 8 best are the fine-tune targets (ids 0β7). Speaker-balanced sampling | |
| β βhours (1β4Γ), per-speaker validation split of 2 %. | |
| **Question labels.** Audiobook readers read most questions with a falling contour, so a raw "?" would teach "slightly | |
| less fall". For clips ending in "?", the final-0.9 s pitch slope decides: rising β keep "?" and oversample 4Γ; | |
| falling but starting with a wh-word (co, kto, gdzie, kiedy, dlaczego, jak, ile, ktΓ³ryβ¦) β keep "?" (falling is correct | |
| Polish); falling otherwise β relabel "?" as "." so the label matches the audio. | |
| **Style tokens.** 18 designed tokens = pitch range {flat, mid, wide} Γ speaking rate {slow, normal, fast} Γ question | |
| {no, yes}, assigned per clip from the reader-relative F0 spread, rate and the question flag (`style_map.json`; the | |
| neutral token is the one the map names `neutral`). Written as the 4th filelist column. | |
| Outputs: `data_v2/mix_v3` (base: 15 WL readers + 5 AZON speakers, 9.5k clips, β 27 h) and `data_v2/mix_v3_target` | |
| (8 readers, 5.2k clips, β 14 h). The filelists are reproducible from the public sources with the commands above. | |
| ## 3. Acoustic model (Matcha-TTS) | |
| Common settings: batch 64, bf16, Adam, `out_size` per Matcha default, checkpoints every N steps plus `final.ckpt`, | |
| validation every 5 epochs, mel statistics computed per dataset (`matcha.utils.generate_data_statistics`). | |
| 0.25 s/step on an H100, 0.31 s/step on an RTX 4090. | |
| | stage | data | init | steps | lr | | |
| |---|---|---|---|---| | |
| | base | `mix_v3` (20 speakers, 18 style tokens) | `matcha_vctk.ckpt`, weights only; speaker, style and symbol tables re-initialised | 40 000 | 1e-4 | | |
| | target | `mix_v3_target` (8 readers) | base `final.ckpt` | 12 000 | 5e-5 | | |
| ```bash | |
| MIX=mix_v3 RUN_NAME=matcha_v3_base STYLE=1 MAX_STEPS=40000 CKPT_EVERY_STEPS=4000 \ | |
| WARM_CKPT=checkpoints/matcha_vctk.ckpt scripts/train_matcha.sh | |
| MIX=mix_v3_target RUN_NAME=matcha_v3_target STYLE=1 MAX_STEPS=12000 CKPT_EVERY_STEPS=2000 \ | |
| WARM_CKPT=runs/matcha_v3_base/checkpoints/final.ckpt EXTRA="model.optimizer.lr=5e-5" scripts/train_matcha.sh | |
| ``` | |
| Warm start is weights-only with a shape-aware partial copy (`scripts/train_matcha_warm.py`): tables that changed size | |
| (speaker, style, symbol embeddings) are copied row by row where shapes overlap, the rest is re-initialised. | |
| Hydra rejects `=` inside override values, so checkpoint files named `step_step=N.ckpt` must be copied to a plain name | |
| before being passed as `+warm_ckpt=`. | |
| The base checkpoint (15 readers + AZON, 40k steps) is not published; to add a new voice, fine-tune the target checkpoint on one to two hours of clean recordings of one speaker with the | |
| target command above, pointing `MIX` at your own filelist. | |
| ## 4. Vocoder (HiFi-GAN fine-tuned to the acoustic model) | |
| Matcha's predicted mels are smoother than real ones (lower harmonic contrast, β 7 dB less energy above 5 kHz). A | |
| vocoder trained only on real mels renders that blur as a phasey second layer and electronic "breaths" in pauses. | |
| Fine-tuning the universal HiFi-GAN on the target model's own mels removes it (the standard "fine-tuning" setting of the | |
| HiFi-GAN paper). | |
| ```bash | |
| # teacher-forced mels for every target clip: MAS alignment on the real mel -> decoder (10 ODE steps, T 0.5), exact GT length | |
| python scripts/gen_mels_tf.py --ckpt runs/matcha_v3_target/checkpoints/final.ckpt \ | |
| --filelist data_v2/mix_v3_target/matcha_train.txt --out vocoder_ft --max-clips 6000 --val 60 --device cuda | |
| # fine-tune generator + discriminators from g_02500000 / do_02500000 | |
| FT_LR=2e-5 FT_DISC_LR_SCALE=0.5 FT_WARMUP=2000 FT_BATCH=16 FT_CKPT_EVERY=5000 \ | |
| scripts/finetune_hifigan.sh vocoder_ft runs/hifigan_ft 93 # 93 epochs Γ 323 steps β 30k steps, 0.2 s/step | |
| ``` | |
| `scripts/patch_hifigan_train.py` adapts the official trainer: it restores the learning rate after loading the optimizer | |
| state from `do_02500000` (which otherwise silently overrides it), trains the generator alone on the mel loss for the | |
| first 2 000 steps before enabling the adversarial and feature-matching losses, and fixes torch β₯ 2 / librosa β₯ 0.10 | |
| incompatibilities. Generator lr 2e-5, discriminators 1e-5, batch 16, segment 8192, 30k steps. The released vocoder is | |
| the 30k-step generator; checkpoints from 15k on sound alike. | |
| ## 5. Evaluation | |
| `scripts/synth_samples.py` (10 conversational sentences, 4β10 ODE steps, temperature 0.5, length scale 0.9) β | |
| `scripts/eval_synth.py`: Whisper large-v3 WER/CER, UTMOS, F0 spread in semitones, speaking rate, silence ratio. | |
| Note that UTMOS does not register the vocoder artefacts described in Β§4 (it even scores the fine-tuned vocoder slightly | |
| lower); the vocoder decision was made by listening, with `scripts/diag_vocoder_copy.py` (copy-synthesis of real clips) | |
| and `scripts/diag_phone_artifacts.py` (per-phoneme roughness) as diagnostics. | |
| ## 6. Inference | |
| - Recommended settings: 4 ODE steps, temperature 0.5, length scale 0.9β1.0, the fine-tuned vocoder, a 10 kHz low-pass | |
| on the output (removes a HiFi-GAN upsampling tone at the Nyquist frequency; `effects.py` in the playground). | |
| - Voice blending: average speaker-embedding rows with weights (`--voice "0.5*6+0.5*3"`). The ONNX exports carry extra | |
| rows with blends already baked in (`scripts/bake_voice.py`), so a blend is just a speaker id. | |
| - Style tokens: an extra embedding added to the speaker embedding; ids and meanings in `style_map.json`. | |
| - ONNX (acoustic model + vocoder in one graph, inputs `x`, `x_lengths`, `scales=[temperature, length_scale]`, `spks`): | |
| `scripts/export_matcha_onnx.py <ckpt> out.onnx --n-timesteps 4 --vocoder-name hifigan_univ_v1 --vocoder-checkpoint-path <fine-tuned g_*>`. | |
| - Measured on an NVIDIA GB10 (PyTorch bf16 + `torch.compile`, batch 1, 4 steps): β 23 ms to first audio, RTF β 0.006. | |
| Apple M-series CPU, ONNX Runtime: β 0.4 s for a 4 s sentence. | |
| ## 7. Licences of the result | |
| Weights trained on CC BY-SA material are released under **CC BY-SA 4.0** (compatible with CC BY-SA 3.0 PL). | |
| `ATTRIBUTION.md` lists every Wolne Lektury book (author, title, reader, director, URL) and the AZON corpus; the warm | |
| start (VCTK) and HiFi-GAN are MIT. Readers' voices are personal attributes not covered by the Creative Commons licence; | |
| the release recommends blended voices under neutral names and disclosure of synthetic speech to listeners. | |