Matcha-TTS-PL / README.md
mcPear's picture
Default synthetic voice = Masiak 1, Staszczyk 0.5, Proszek 0.5 (UTMOS 3.38); female preset with three base-model women at 0.5
1863d68 verified
|
Raw History Blame Contribute Delete
9.25 kB
---
language: pl
license: cc-by-sa-4.0
library_name: matcha-tts
pipeline_tag: text-to-speech
tags: [tts, polish, matcha-tts, flow-matching, hifigan, multi-speaker, style-tokens, robot]
---
# Matcha-TTS-PL — Polish Matcha-TTS for a conversational robot
Non-autoregressive Polish text-to-speech (Matcha-TTS, optimal-transport conditional flow matching, 20.9 M parameters)
built for a humanoid robot: real-time on a small GPU (≈ 23 ms to first audio on an NVIDIA GB10, 4 ODE steps),
8 blendable reader voices, 18 style tokens (pitch range × speaking rate × question), and a **HiFi-GAN vocoder fine-tuned
to this model**, which removes the phasey layer a stock vocoder adds to predicted mels. Trained from a VCTK warm start on
consistency-filtered Wolne Lektury audiobooks (prose only, text/audio agreement verified with Whisper) plus AZON
spontaneous speech.
## Files
| path | what |
|---|---|
| `model/matcha_pl_target.ckpt` | **the released acoustic model**: Lightning checkpoint, 20 speaker rows (ids 0–7 = target readers), 18 style rows |
| `vocoder/hifigan_pl.pt` | HiFi-GAN generator fine-tuned on this model's mels (**use this one**); `vocoder/g_02500000_universal` = the stock universal vocoder for comparison |
| `onnx/matcha_pl_t2.onnx`, `onnx/matcha_pl_t4.onnx` | acoustic model + fine-tuned vocoder in one ONNX graph, 2 / 4 ODE steps; inputs `x` (phoneme ids), `x_lengths`, `scales=[temperature, length_scale]`, `spk_emb` (float32 [1, 64]); outputs `wav`, `wav_lengths` |
| `onnx/voices.json` | speaker and style embedding tables for building `spk_emb` (any blend, any style), plus five synthetic voice presets (`synthetic` is the default) |
| `data/speaker_map.json`, `data/speakers.json` | speaker id → reader |
| `data/style_map.json`, `data/style_centroids.json` | style token definitions |
| `samples/` | synthesised test sentences (`manifest.csv`: file, voice, text) |
| `RECIPE.md` | the full training procedure (data, base, target, vocoder fine-tune) |
| `ATTRIBUTION.md`, `LICENSE` | data attribution (every book, reader, director) and CC BY-SA 4.0 |
| (external) | UI, CLI tools and the Matcha-TTS patch the checkpoint needs: [github.com/machinekind/tts-pl-playground](https://github.com/machinekind/tts-pl-playground) |
## Quick start
```bash
git clone https://github.com/machinekind/tts-pl-playground.git && cd tts-pl-playground # follow its README (Matcha-TTS clone + patch)
python -c "from huggingface_hub import snapshot_download as d; d('machinekind/Matcha-TTS-PL', local_dir='.')"
python playground.py --port 8771 # UI: voices, blends, style tokens, effects, sequence composer, video visualizer
python synth_samples.py --ckpt model/matcha_pl_target.ckpt --vocoder vocoder/hifigan_pl.pt \
--sentences my_sentences.txt --voice "1*2+0.5*4+0.5*5" --steps 4 --temperature 0 --out out/
```
ONNX Runtime (CPU, GPU or in the browser), no PyTorch needed at run time:
```python
import json, onnxruntime as ort, numpy as np
vj = json.load(open("onnx/voices.json"))
E = lambda i: np.array(vj["speakers"][str(i)]["emb"], np.float32)
w = {int(i): x for i, x in vj["presets"]["synthetic"]["weights"].items() if x > 0} # the default synthetic voice
v = sum(x * E(i) for i, x in w.items()) / sum(w.values())
spk_emb = v * sum(x * np.linalg.norm(E(i)) for i, x in w.items()) / sum(w.values()) / np.linalg.norm(v) # rescale: a plain average is too short
# add vj["styles"]["14"]["emb"] for a wider pitch range
sess = ort.InferenceSession("onnx/matcha_pl_t4.onnx")
x = phonemes # int64 [1, T]: espeak-ng "pl" ids from the playground's polish_cleaners (see RECIPE.md §1)
wav, n = sess.run(None, {"x": x, "x_lengths": np.array([x.shape[1]]), "scales": np.array([0.0, 0.95], np.float32), "spk_emb": spk_emb[None]})
```
## Voices
Speaker ids 0–7 are the target readers (0 Bartosz Bielenia, 1 Katarzyna Faszczewska, 2 Wojciech Masiak, 3 Bartosz Głogowski, 4 Jan Staszczyk, 5 Marek Proszek, 6 Piotr Kopa, 7 Radosław Krzyżowski); ids 8–19 are base-training speakers kept for completeness.
Any voice is a weighted blend of speaker embeddings, optionally plus a style embedding:
`spk_emb = Σ w_i · E_spk[i] / Σ w_i (+ E_style[k])`. In PyTorch use `--voice "0.33*1+0.33*2+0.34*0"` and `--style k`;
the ONNX graphs take `spk_emb` directly, with both tables in `onnx/voices.json`, so blends and styles need no re-export.
The browser playground exposes this as a mixer.
`onnx/voices.json` ships five synthetic voice presets (`synthetic` = the default, `synthetic2`–`synthetic5`), each a weighted blend of readers. Rescale every
blend to the average length of its component vectors: a plain average is much shorter than any trained voice and makes
short sentences unstable. Blends of
several readers are the recommended way to deploy (see *Licence and attribution* on voice rights).
## Style tokens
An extra embedding added to the speaker embedding: pass `styles=<id>` to `synthesise` in PyTorch; in ONNX add `styles[id].emb` from `onnx/voices.json` to `spk_emb`.
Neutral = `7`. Pitch-range tokens move the spread by about ±1.3 semitones; the question tokens give a rising
terminal contour for yes/no questions (Polish wh-questions fall — write real punctuation, "?" drives intonation).
| id | label |
|---|---|
| 0 | flat-range · slow |
| 1 | flat-range · slow · question |
| 2 | flat-range · normal |
| 3 | flat-range · normal · question |
| 4 | flat-range · fast |
| 5 | flat-range · fast · question |
| 6 | mid-range · slow |
| 7 | mid-range · slow · question |
| 8 | mid-range · normal |
| 9 | mid-range · normal · question |
| 10 | mid-range · fast |
| 11 | mid-range · fast · question |
| 12 | wide-range · slow |
| 13 | wide-range · slow · question |
| 14 | wide-range · normal |
| 15 | wide-range · normal · question |
| 16 | wide-range · fast |
| 17 | wide-range · fast · question |
## Quality
### 10 conversational test sentences per voice (4 ODE steps, T 0.5, fine-tuned vocoder; mix0 = the default synthetic voice)
| voice | n | WER | CER | UTMOS | F0 spread [st] | chars/s | silence % |
|---|---|---|---|---|---|---|---|
| Bartosz Bielenia | 10 | 0.012 | 0.002 | 3.21 | 2.59 | 10.5 | 21 |
| Katarzyna Faszczewska | 10 | 0.012 | 0.002 | 3.35 | 2.99 | 10.1 | 21 |
| Wojciech Masiak | 10 | 0.023 | 0.006 | 3.30 | 3.29 | 11.7 | 14 |
| Bartosz Głogowski | 10 | 0.000 | 0.000 | 3.26 | 4.10 | 10.9 | 17 |
| Jan Staszczyk | 10 | 0.047 | 0.061 | 3.07 | 3.54 | 10.7 | 15 |
| Marek Proszek | 10 | 0.047 | 0.063 | 3.19 | 3.73 | 11.0 | 16 |
| Piotr Kopa | 10 | 0.058 | 0.069 | 3.27 | 2.61 | 10.0 | 11 |
| Radosław Krzyżowski | 10 | 0.035 | 0.015 | 2.81 | 2.04 | 10.2 | 22 |
| mix0 | 10 | 0.012 | 0.002 | 3.38 | 2.84 | 11.5 | 15 |
Whisper large-v3 WER/CER, UTMOS (`tarepan/SpeechMOS`), pitch spread. UTMOS does not capture the vocoder artefacts the
fine-tune removes; the vocoder choice was made by listening (see `RECIPE.md` §4–5).
Latency: NVIDIA GB10, PyTorch bf16 + `torch.compile`, batch 1, 4 steps ≈ 23 ms to first audio, RTF ≈ 0.006.
Apple M-series CPU, ONNX Runtime, 4 steps ≈ 0.4 s for a 4 s sentence. Recommended runtime settings: 4 ODE steps,
temperature 0 (deterministic, most consistent across sentences) up to 0.8 (the playground default: livelier, more variation), length scale 0.9–1.0.
## Known limitations
Phrase endings are flatter than a human reader's (the model averages final contours); short one-word replies are less
natural than full sentences; only readers' prose style is covered (no shouting, whispering or singing).
## Training procedure (summary)
`RECIPE.md` has the complete, reproducible version. Data: Wolne Lektury prose + AZON, per-clip UTMOS/DNSMOS/F0 statistics,
Whisper CER filter, reader ranking by consistency (top 15 → base, top 8 → target), question labels corrected from the
measured final pitch, 18 designed style tokens. Base: 40k steps from `matcha_vctk` (batch 64, bf16, lr 1e-4).
Target: 12k steps on the 8 readers (lr 5e-5). Vocoder: HiFi-GAN universal fine-tuned 30k steps on the target model's
teacher-forced mels (generator lr 2e-5, discriminators 1e-5, 2k-step generator-only warm-up). About 6 GPU-hours on an RTX 4090.
## Licence and attribution
- **Weights: CC BY-SA 4.0** (`LICENSE`). The training audio is CC BY-SA 3.0 PL (Wolne Lektury) and CC BY-SA 4.0 (AZON);
ShareAlike propagates to the weights. Every book, reader and director is listed in `ATTRIBUTION.md` — keep that file
with any redistribution or derivative.
- Warm start: Matcha-TTS `matcha_vctk.ckpt` (MIT; VCTK corpus CC BY 4.0, CSTR, University of Edinburgh). Vocoder: HiFi-GAN
universal v1 (MIT), fine-tuned here. Code: Matcha-TTS (MIT), jik876/hifi-gan (MIT), espeak-ng (GPL-3.0, runtime dependency).
- **Voices are personal attributes.** The CC licence covers the recordings, not the readers' personality rights. The
recommended deployment is a blend of two or more readers under a neutral voice name (the default synthetic voice is one);
using a single reader's voice commercially should be cleared with the reader or Wolne Lektury. Readers are named here
only as data sources.
- Synthetic speech should be disclosed as such where the listener could otherwise take it for a person (EU AI Act, art. 50).