|
Download README.md from machinekind/Matcha-TTS-PL: direct link, hf CLI and curl.
- Browser
- Download file 9.25 kB
-
https://huggingface.co/machinekind/Matcha-TTS-PL/resolve/main/README.md
- Command line
-
hf download hf://machinekind/Matcha-TTS-PL/README.md
-
curl -L -o README.md https://huggingface.co/machinekind/Matcha-TTS-PL/resolve/main/README.md
9.25 kB
| language: pl | |
| license: cc-by-sa-4.0 | |
| library_name: matcha-tts | |
| pipeline_tag: text-to-speech | |
| tags: [tts, polish, matcha-tts, flow-matching, hifigan, multi-speaker, style-tokens, robot] | |
| # Matcha-TTS-PL — Polish Matcha-TTS for a conversational robot | |
| Non-autoregressive Polish text-to-speech (Matcha-TTS, optimal-transport conditional flow matching, 20.9 M parameters) | |
| built for a humanoid robot: real-time on a small GPU (≈ 23 ms to first audio on an NVIDIA GB10, 4 ODE steps), | |
| 8 blendable reader voices, 18 style tokens (pitch range × speaking rate × question), and a **HiFi-GAN vocoder fine-tuned | |
| to this model**, which removes the phasey layer a stock vocoder adds to predicted mels. Trained from a VCTK warm start on | |
| consistency-filtered Wolne Lektury audiobooks (prose only, text/audio agreement verified with Whisper) plus AZON | |
| spontaneous speech. | |
| ## Files | |
| | path | what | | |
| |---|---| | |
| | `model/matcha_pl_target.ckpt` | **the released acoustic model**: Lightning checkpoint, 20 speaker rows (ids 0–7 = target readers), 18 style rows | | |
| | `vocoder/hifigan_pl.pt` | HiFi-GAN generator fine-tuned on this model's mels (**use this one**); `vocoder/g_02500000_universal` = the stock universal vocoder for comparison | | |
| | `onnx/matcha_pl_t2.onnx`, `onnx/matcha_pl_t4.onnx` | acoustic model + fine-tuned vocoder in one ONNX graph, 2 / 4 ODE steps; inputs `x` (phoneme ids), `x_lengths`, `scales=[temperature, length_scale]`, `spk_emb` (float32 [1, 64]); outputs `wav`, `wav_lengths` | | |
| | `onnx/voices.json` | speaker and style embedding tables for building `spk_emb` (any blend, any style), plus five synthetic voice presets (`synthetic` is the default) | | |
| | `data/speaker_map.json`, `data/speakers.json` | speaker id → reader | | |
| | `data/style_map.json`, `data/style_centroids.json` | style token definitions | | |
| | `samples/` | synthesised test sentences (`manifest.csv`: file, voice, text) | | |
| | `RECIPE.md` | the full training procedure (data, base, target, vocoder fine-tune) | | |
| | `ATTRIBUTION.md`, `LICENSE` | data attribution (every book, reader, director) and CC BY-SA 4.0 | | |
| | (external) | UI, CLI tools and the Matcha-TTS patch the checkpoint needs: [github.com/machinekind/tts-pl-playground](https://github.com/machinekind/tts-pl-playground) | | |
| ## Quick start | |
| ```bash | |
| git clone https://github.com/machinekind/tts-pl-playground.git && cd tts-pl-playground # follow its README (Matcha-TTS clone + patch) | |
| python -c "from huggingface_hub import snapshot_download as d; d('machinekind/Matcha-TTS-PL', local_dir='.')" | |
| python playground.py --port 8771 # UI: voices, blends, style tokens, effects, sequence composer, video visualizer | |
| python synth_samples.py --ckpt model/matcha_pl_target.ckpt --vocoder vocoder/hifigan_pl.pt \ | |
| --sentences my_sentences.txt --voice "1*2+0.5*4+0.5*5" --steps 4 --temperature 0 --out out/ | |
| ``` | |
| ONNX Runtime (CPU, GPU or in the browser), no PyTorch needed at run time: | |
| ```python | |
| import json, onnxruntime as ort, numpy as np | |
| vj = json.load(open("onnx/voices.json")) | |
| E = lambda i: np.array(vj["speakers"][str(i)]["emb"], np.float32) | |
| w = {int(i): x for i, x in vj["presets"]["synthetic"]["weights"].items() if x > 0} # the default synthetic voice | |
| v = sum(x * E(i) for i, x in w.items()) / sum(w.values()) | |
| spk_emb = v * sum(x * np.linalg.norm(E(i)) for i, x in w.items()) / sum(w.values()) / np.linalg.norm(v) # rescale: a plain average is too short | |
| # add vj["styles"]["14"]["emb"] for a wider pitch range | |
| sess = ort.InferenceSession("onnx/matcha_pl_t4.onnx") | |
| x = phonemes # int64 [1, T]: espeak-ng "pl" ids from the playground's polish_cleaners (see RECIPE.md §1) | |
| wav, n = sess.run(None, {"x": x, "x_lengths": np.array([x.shape[1]]), "scales": np.array([0.0, 0.95], np.float32), "spk_emb": spk_emb[None]}) | |
| ``` | |
| ## Voices | |
| Speaker ids 0–7 are the target readers (0 Bartosz Bielenia, 1 Katarzyna Faszczewska, 2 Wojciech Masiak, 3 Bartosz Głogowski, 4 Jan Staszczyk, 5 Marek Proszek, 6 Piotr Kopa, 7 Radosław Krzyżowski); ids 8–19 are base-training speakers kept for completeness. | |
| Any voice is a weighted blend of speaker embeddings, optionally plus a style embedding: | |
| `spk_emb = Σ w_i · E_spk[i] / Σ w_i (+ E_style[k])`. In PyTorch use `--voice "0.33*1+0.33*2+0.34*0"` and `--style k`; | |
| the ONNX graphs take `spk_emb` directly, with both tables in `onnx/voices.json`, so blends and styles need no re-export. | |
| The browser playground exposes this as a mixer. | |
| `onnx/voices.json` ships five synthetic voice presets (`synthetic` = the default, `synthetic2`–`synthetic5`), each a weighted blend of readers. Rescale every | |
| blend to the average length of its component vectors: a plain average is much shorter than any trained voice and makes | |
| short sentences unstable. Blends of | |
| several readers are the recommended way to deploy (see *Licence and attribution* on voice rights). | |
| ## Style tokens | |
| An extra embedding added to the speaker embedding: pass `styles=<id>` to `synthesise` in PyTorch; in ONNX add `styles[id].emb` from `onnx/voices.json` to `spk_emb`. | |
| Neutral = `7`. Pitch-range tokens move the spread by about ±1.3 semitones; the question tokens give a rising | |
| terminal contour for yes/no questions (Polish wh-questions fall — write real punctuation, "?" drives intonation). | |
| | id | label | | |
| |---|---| | |
| | 0 | flat-range · slow | | |
| | 1 | flat-range · slow · question | | |
| | 2 | flat-range · normal | | |
| | 3 | flat-range · normal · question | | |
| | 4 | flat-range · fast | | |
| | 5 | flat-range · fast · question | | |
| | 6 | mid-range · slow | | |
| | 7 | mid-range · slow · question | | |
| | 8 | mid-range · normal | | |
| | 9 | mid-range · normal · question | | |
| | 10 | mid-range · fast | | |
| | 11 | mid-range · fast · question | | |
| | 12 | wide-range · slow | | |
| | 13 | wide-range · slow · question | | |
| | 14 | wide-range · normal | | |
| | 15 | wide-range · normal · question | | |
| | 16 | wide-range · fast | | |
| | 17 | wide-range · fast · question | | |
| ## Quality | |
| ### 10 conversational test sentences per voice (4 ODE steps, T 0.5, fine-tuned vocoder; mix0 = the default synthetic voice) | |
| | voice | n | WER | CER | UTMOS | F0 spread [st] | chars/s | silence % | | |
| |---|---|---|---|---|---|---|---| | |
| | Bartosz Bielenia | 10 | 0.012 | 0.002 | 3.21 | 2.59 | 10.5 | 21 | | |
| | Katarzyna Faszczewska | 10 | 0.012 | 0.002 | 3.35 | 2.99 | 10.1 | 21 | | |
| | Wojciech Masiak | 10 | 0.023 | 0.006 | 3.30 | 3.29 | 11.7 | 14 | | |
| | Bartosz Głogowski | 10 | 0.000 | 0.000 | 3.26 | 4.10 | 10.9 | 17 | | |
| | Jan Staszczyk | 10 | 0.047 | 0.061 | 3.07 | 3.54 | 10.7 | 15 | | |
| | Marek Proszek | 10 | 0.047 | 0.063 | 3.19 | 3.73 | 11.0 | 16 | | |
| | Piotr Kopa | 10 | 0.058 | 0.069 | 3.27 | 2.61 | 10.0 | 11 | | |
| | Radosław Krzyżowski | 10 | 0.035 | 0.015 | 2.81 | 2.04 | 10.2 | 22 | | |
| | mix0 | 10 | 0.012 | 0.002 | 3.38 | 2.84 | 11.5 | 15 | | |
| Whisper large-v3 WER/CER, UTMOS (`tarepan/SpeechMOS`), pitch spread. UTMOS does not capture the vocoder artefacts the | |
| fine-tune removes; the vocoder choice was made by listening (see `RECIPE.md` §4–5). | |
| Latency: NVIDIA GB10, PyTorch bf16 + `torch.compile`, batch 1, 4 steps ≈ 23 ms to first audio, RTF ≈ 0.006. | |
| Apple M-series CPU, ONNX Runtime, 4 steps ≈ 0.4 s for a 4 s sentence. Recommended runtime settings: 4 ODE steps, | |
| temperature 0 (deterministic, most consistent across sentences) up to 0.8 (the playground default: livelier, more variation), length scale 0.9–1.0. | |
| ## Known limitations | |
| Phrase endings are flatter than a human reader's (the model averages final contours); short one-word replies are less | |
| natural than full sentences; only readers' prose style is covered (no shouting, whispering or singing). | |
| ## Training procedure (summary) | |
| `RECIPE.md` has the complete, reproducible version. Data: Wolne Lektury prose + AZON, per-clip UTMOS/DNSMOS/F0 statistics, | |
| Whisper CER filter, reader ranking by consistency (top 15 → base, top 8 → target), question labels corrected from the | |
| measured final pitch, 18 designed style tokens. Base: 40k steps from `matcha_vctk` (batch 64, bf16, lr 1e-4). | |
| Target: 12k steps on the 8 readers (lr 5e-5). Vocoder: HiFi-GAN universal fine-tuned 30k steps on the target model's | |
| teacher-forced mels (generator lr 2e-5, discriminators 1e-5, 2k-step generator-only warm-up). About 6 GPU-hours on an RTX 4090. | |
| ## Licence and attribution | |
| - **Weights: CC BY-SA 4.0** (`LICENSE`). The training audio is CC BY-SA 3.0 PL (Wolne Lektury) and CC BY-SA 4.0 (AZON); | |
| ShareAlike propagates to the weights. Every book, reader and director is listed in `ATTRIBUTION.md` — keep that file | |
| with any redistribution or derivative. | |
| - Warm start: Matcha-TTS `matcha_vctk.ckpt` (MIT; VCTK corpus CC BY 4.0, CSTR, University of Edinburgh). Vocoder: HiFi-GAN | |
| universal v1 (MIT), fine-tuned here. Code: Matcha-TTS (MIT), jik876/hifi-gan (MIT), espeak-ng (GPL-3.0, runtime dependency). | |
| - **Voices are personal attributes.** The CC licence covers the recordings, not the readers' personality rights. The | |
| recommended deployment is a blend of two or more readers under a neutral voice name (the default synthetic voice is one); | |
| using a single reader's voice commercially should be cleared with the reader or Wolne Lektury. Readers are named here | |
| only as data sources. | |
| - Synthetic speech should be disclosed as such where the listener could otherwise take it for a person (EU AI Act, art. 50). | |