File size: 9,251 Bytes
c24c082 29857f3 806857e c24c082 1863d68 c24c082 29857f3 c24c082 29857f3 934d64f a328640 29857f3 c24c082 4153160 c24c082 29857f3 806857e a328640 29857f3 c24c082 29857f3 c24c082 934d64f c24c082 29857f3 1863d68 c24c082 f85e86b c24c082 934d64f c24c082 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 | ---
language: pl
license: cc-by-sa-4.0
library_name: matcha-tts
pipeline_tag: text-to-speech
tags: [tts, polish, matcha-tts, flow-matching, hifigan, multi-speaker, style-tokens, robot]
---
# Matcha-TTS-PL — Polish Matcha-TTS for a conversational robot
Non-autoregressive Polish text-to-speech (Matcha-TTS, optimal-transport conditional flow matching, 20.9 M parameters)
built for a humanoid robot: real-time on a small GPU (≈ 23 ms to first audio on an NVIDIA GB10, 4 ODE steps),
8 blendable reader voices, 18 style tokens (pitch range × speaking rate × question), and a **HiFi-GAN vocoder fine-tuned
to this model**, which removes the phasey layer a stock vocoder adds to predicted mels. Trained from a VCTK warm start on
consistency-filtered Wolne Lektury audiobooks (prose only, text/audio agreement verified with Whisper) plus AZON
spontaneous speech.
## Files
| path | what |
|---|---|
| `model/matcha_pl_target.ckpt` | **the released acoustic model**: Lightning checkpoint, 20 speaker rows (ids 0–7 = target readers), 18 style rows |
| `vocoder/hifigan_pl.pt` | HiFi-GAN generator fine-tuned on this model's mels (**use this one**); `vocoder/g_02500000_universal` = the stock universal vocoder for comparison |
| `onnx/matcha_pl_t2.onnx`, `onnx/matcha_pl_t4.onnx` | acoustic model + fine-tuned vocoder in one ONNX graph, 2 / 4 ODE steps; inputs `x` (phoneme ids), `x_lengths`, `scales=[temperature, length_scale]`, `spk_emb` (float32 [1, 64]); outputs `wav`, `wav_lengths` |
| `onnx/voices.json` | speaker and style embedding tables for building `spk_emb` (any blend, any style), plus five synthetic voice presets (`synthetic` is the default) |
| `data/speaker_map.json`, `data/speakers.json` | speaker id → reader |
| `data/style_map.json`, `data/style_centroids.json` | style token definitions |
| `samples/` | synthesised test sentences (`manifest.csv`: file, voice, text) |
| `RECIPE.md` | the full training procedure (data, base, target, vocoder fine-tune) |
| `ATTRIBUTION.md`, `LICENSE` | data attribution (every book, reader, director) and CC BY-SA 4.0 |
| (external) | UI, CLI tools and the Matcha-TTS patch the checkpoint needs: [github.com/machinekind/tts-pl-playground](https://github.com/machinekind/tts-pl-playground) |
## Quick start
```bash
git clone https://github.com/machinekind/tts-pl-playground.git && cd tts-pl-playground # follow its README (Matcha-TTS clone + patch)
python -c "from huggingface_hub import snapshot_download as d; d('machinekind/Matcha-TTS-PL', local_dir='.')"
python playground.py --port 8771 # UI: voices, blends, style tokens, effects, sequence composer, video visualizer
python synth_samples.py --ckpt model/matcha_pl_target.ckpt --vocoder vocoder/hifigan_pl.pt \
--sentences my_sentences.txt --voice "1*2+0.5*4+0.5*5" --steps 4 --temperature 0 --out out/
```
ONNX Runtime (CPU, GPU or in the browser), no PyTorch needed at run time:
```python
import json, onnxruntime as ort, numpy as np
vj = json.load(open("onnx/voices.json"))
E = lambda i: np.array(vj["speakers"][str(i)]["emb"], np.float32)
w = {int(i): x for i, x in vj["presets"]["synthetic"]["weights"].items() if x > 0} # the default synthetic voice
v = sum(x * E(i) for i, x in w.items()) / sum(w.values())
spk_emb = v * sum(x * np.linalg.norm(E(i)) for i, x in w.items()) / sum(w.values()) / np.linalg.norm(v) # rescale: a plain average is too short
# add vj["styles"]["14"]["emb"] for a wider pitch range
sess = ort.InferenceSession("onnx/matcha_pl_t4.onnx")
x = phonemes # int64 [1, T]: espeak-ng "pl" ids from the playground's polish_cleaners (see RECIPE.md §1)
wav, n = sess.run(None, {"x": x, "x_lengths": np.array([x.shape[1]]), "scales": np.array([0.0, 0.95], np.float32), "spk_emb": spk_emb[None]})
```
## Voices
Speaker ids 0–7 are the target readers (0 Bartosz Bielenia, 1 Katarzyna Faszczewska, 2 Wojciech Masiak, 3 Bartosz Głogowski, 4 Jan Staszczyk, 5 Marek Proszek, 6 Piotr Kopa, 7 Radosław Krzyżowski); ids 8–19 are base-training speakers kept for completeness.
Any voice is a weighted blend of speaker embeddings, optionally plus a style embedding:
`spk_emb = Σ w_i · E_spk[i] / Σ w_i (+ E_style[k])`. In PyTorch use `--voice "0.33*1+0.33*2+0.34*0"` and `--style k`;
the ONNX graphs take `spk_emb` directly, with both tables in `onnx/voices.json`, so blends and styles need no re-export.
The browser playground exposes this as a mixer.
`onnx/voices.json` ships five synthetic voice presets (`synthetic` = the default, `synthetic2`–`synthetic5`), each a weighted blend of readers. Rescale every
blend to the average length of its component vectors: a plain average is much shorter than any trained voice and makes
short sentences unstable. Blends of
several readers are the recommended way to deploy (see *Licence and attribution* on voice rights).
## Style tokens
An extra embedding added to the speaker embedding: pass `styles=<id>` to `synthesise` in PyTorch; in ONNX add `styles[id].emb` from `onnx/voices.json` to `spk_emb`.
Neutral = `7`. Pitch-range tokens move the spread by about ±1.3 semitones; the question tokens give a rising
terminal contour for yes/no questions (Polish wh-questions fall — write real punctuation, "?" drives intonation).
| id | label |
|---|---|
| 0 | flat-range · slow |
| 1 | flat-range · slow · question |
| 2 | flat-range · normal |
| 3 | flat-range · normal · question |
| 4 | flat-range · fast |
| 5 | flat-range · fast · question |
| 6 | mid-range · slow |
| 7 | mid-range · slow · question |
| 8 | mid-range · normal |
| 9 | mid-range · normal · question |
| 10 | mid-range · fast |
| 11 | mid-range · fast · question |
| 12 | wide-range · slow |
| 13 | wide-range · slow · question |
| 14 | wide-range · normal |
| 15 | wide-range · normal · question |
| 16 | wide-range · fast |
| 17 | wide-range · fast · question |
## Quality
### 10 conversational test sentences per voice (4 ODE steps, T 0.5, fine-tuned vocoder; mix0 = the default synthetic voice)
| voice | n | WER | CER | UTMOS | F0 spread [st] | chars/s | silence % |
|---|---|---|---|---|---|---|---|
| Bartosz Bielenia | 10 | 0.012 | 0.002 | 3.21 | 2.59 | 10.5 | 21 |
| Katarzyna Faszczewska | 10 | 0.012 | 0.002 | 3.35 | 2.99 | 10.1 | 21 |
| Wojciech Masiak | 10 | 0.023 | 0.006 | 3.30 | 3.29 | 11.7 | 14 |
| Bartosz Głogowski | 10 | 0.000 | 0.000 | 3.26 | 4.10 | 10.9 | 17 |
| Jan Staszczyk | 10 | 0.047 | 0.061 | 3.07 | 3.54 | 10.7 | 15 |
| Marek Proszek | 10 | 0.047 | 0.063 | 3.19 | 3.73 | 11.0 | 16 |
| Piotr Kopa | 10 | 0.058 | 0.069 | 3.27 | 2.61 | 10.0 | 11 |
| Radosław Krzyżowski | 10 | 0.035 | 0.015 | 2.81 | 2.04 | 10.2 | 22 |
| mix0 | 10 | 0.012 | 0.002 | 3.38 | 2.84 | 11.5 | 15 |
Whisper large-v3 WER/CER, UTMOS (`tarepan/SpeechMOS`), pitch spread. UTMOS does not capture the vocoder artefacts the
fine-tune removes; the vocoder choice was made by listening (see `RECIPE.md` §4–5).
Latency: NVIDIA GB10, PyTorch bf16 + `torch.compile`, batch 1, 4 steps ≈ 23 ms to first audio, RTF ≈ 0.006.
Apple M-series CPU, ONNX Runtime, 4 steps ≈ 0.4 s for a 4 s sentence. Recommended runtime settings: 4 ODE steps,
temperature 0 (deterministic, most consistent across sentences) up to 0.8 (the playground default: livelier, more variation), length scale 0.9–1.0.
## Known limitations
Phrase endings are flatter than a human reader's (the model averages final contours); short one-word replies are less
natural than full sentences; only readers' prose style is covered (no shouting, whispering or singing).
## Training procedure (summary)
`RECIPE.md` has the complete, reproducible version. Data: Wolne Lektury prose + AZON, per-clip UTMOS/DNSMOS/F0 statistics,
Whisper CER filter, reader ranking by consistency (top 15 → base, top 8 → target), question labels corrected from the
measured final pitch, 18 designed style tokens. Base: 40k steps from `matcha_vctk` (batch 64, bf16, lr 1e-4).
Target: 12k steps on the 8 readers (lr 5e-5). Vocoder: HiFi-GAN universal fine-tuned 30k steps on the target model's
teacher-forced mels (generator lr 2e-5, discriminators 1e-5, 2k-step generator-only warm-up). About 6 GPU-hours on an RTX 4090.
## Licence and attribution
- **Weights: CC BY-SA 4.0** (`LICENSE`). The training audio is CC BY-SA 3.0 PL (Wolne Lektury) and CC BY-SA 4.0 (AZON);
ShareAlike propagates to the weights. Every book, reader and director is listed in `ATTRIBUTION.md` — keep that file
with any redistribution or derivative.
- Warm start: Matcha-TTS `matcha_vctk.ckpt` (MIT; VCTK corpus CC BY 4.0, CSTR, University of Edinburgh). Vocoder: HiFi-GAN
universal v1 (MIT), fine-tuned here. Code: Matcha-TTS (MIT), jik876/hifi-gan (MIT), espeak-ng (GPL-3.0, runtime dependency).
- **Voices are personal attributes.** The CC licence covers the recordings, not the readers' personality rights. The
recommended deployment is a blend of two or more readers under a neutral voice name (the default synthetic voice is one);
using a single reader's voice commercially should be cleared with the reader or Wolne Lektury. Readers are named here
only as data sources.
- Synthetic speech should be disclosed as such where the listener could otherwise take it for a person (EU AI Act, art. 50).
|