Matcha-TTS-PL v2 — one synthetic voice, modern Polish, English words
Second version of Matcha-TTS-PL: the same small, fast, non-autoregressive Matcha-TTS (20.9 M parameters, real-time on a small GPU or in a browser) distilled from VoxCPM2, a 2B-parameter TTS, onto a single synthetic voice. What changes compared with v1:
- One original voice,
fav2— designed from a text description with VoxCPM2 (OpenBMB, Apache-2.0) and frozen as a prompt cache; not a recording or a clone of a real person (closest of our reference readers: speaker similarity 0.83). - Modern Polish of a humanoid assistant — 13.1 h of training speech: assistant speech in 18 everyday domains, modern corpora (NKJP, Wikinews, KPWr, ParlaMint, OpenAssistant, Wikibooks, Wikivoyage) and reviewed YouTube transcripts; two thirds of it in multi-sentence passages (10–25 s), so prosody is planned across sentences.
- English words inside Polish sentences — "Po meetingu wyślę ci feedback na maila" is read with English phonemes:
English words are detected automatically and phonemized by espeak-ng en-us, the rest by espeak-ng pl
(
english/). No pronunciation dictionary needed for common loanwords and brands. - A HiFi-GAN fine-tuned on this voice, with real ground truth (the teacher's audio) — the v1 vocoder was tuned for the audiobook readers and blurs this voice.
Try it in the browser: spaces/machinekind/matcha-tts-pl-v2.
Quality (12 held-out sentences, Mac CPU, 4 ODE steps, temperature 0.8)
| system | UTMOS ↑ | Whisper WER ↓ | similarity to fav2 ↑ |
|---|---|---|---|
| VoxCPM2 teacher (2B, autoregressive) | 3.68 | 1.3 % | 0.883 |
| v2 (this model + its vocoder) | 3.40 | 1.3 % | 0.850 |
| v1 model with the new voice slot, before fine-tuning | 2.91 | 2.7 % | 0.725 |
WER counts Whisper's spelling too: the teacher's two "errors" are "krutki" for krótki and "HWSR" for the robot's name.
Speaking rate: the released default tempo is 0.8 (length_scale); at 1.0 v2 speaks about 19 % slower than the teacher.
Files
| path | what |
|---|---|
onnx/matcha_pl_t4.onnx, onnx/matcha_pl_t2.onnx |
acoustic model + fav2 vocoder in one graph, 4 / 2 ODE steps; inputs x (phoneme ids), x_lengths, scales=[temperature, length_scale], spk_emb (float32 [1, 64]); outputs wav (22.05 kHz), wav_lengths |
onnx/voices.json |
the fav2 speaker embedding, 18 style embeddings, defaults |
model/matcha_pl_v2.ckpt |
PyTorch checkpoint (Matcha-TTS + the tts-pl patch: style tokens, Polish cleaners) — speaker 20 = fav2 |
vocoder/generator_fav2, vocoder/config.json |
HiFi-GAN (v1 architecture) fine-tuned on this model's mels ↔ the fav2 audio |
english/ |
English detection (english_detect.py, english_words.txt; browser: english.js, english_data.json) and the mixed phonemizer (polish_mixed_cleaners.py; browser: cleaner.js) |
samples/ |
the 12 held-out sentences, v2 and teacher |
data/ |
style token map, speaker map |
ATTRIBUTION.md |
every text source of the training data, and the base model's data |
Quick start (ONNX, Python)
import sys, json, numpy as np, onnxruntime as ort, soundfile as sf
sys.path.append("english") # english_detect.py + english_words.txt (optional fallback: pip install wordfreq)
from english_detect import mark_english
from matcha.text import text_to_sequence # Matcha-TTS with the tts-pl cleaners: paste english/polish_cleaners.py and
from matcha.utils.utils import intersperse # english/polish_mixed_cleaners.py into matcha/text/cleaners.py
text = mark_english("Po meetingu wyślę ci feedback na maila, a w weekend zrobię update.")
ids = np.array(intersperse(text_to_sequence(text, ["polish_mixed_cleaners"])[0], 0), np.int64)
vj = json.load(open("onnx/voices.json"))
spk = np.array(vj["speakers"]["0"]["emb"], np.float32) + np.array(vj["styles"]["7"]["emb"], np.float32) # 7 = neutral style
sess = ort.InferenceSession("onnx/matcha_pl_t4.onnx")
wav, n = sess.run(None, {"x": ids[None], "x_lengths": np.array([ids.size]), "scales": np.array([0.8, 0.8], np.float32), "spk_emb": spk[None]})
sf.write("out.wav", wav[0, :n[0]], 22050) # scales = [temperature, length_scale (tempo)]
Text front end: mark_english() wraps English words in braces (Po {meeting}u wyślę ci {feedback}.), the cleaner
phonemizes braced spans with espeak-ng en-us and the rest with pl (with the Polish ending of an inflected English
word glued on unstressed, and one-letter prepositions before English words handled by rule). You can brace any word
yourself. The browser port produces the training phonemes character for character (tested on 300 sentences).
Training
- Base: Matcha-TTS-PL v1 (target model), new speaker row initialised from the closest real reader; same mel statistics, style tokens and phoneme table (all en-us phones were already in the symbol table, now trained).
- Data: 6 493 clips, 13.1 h, generated by VoxCPM2 with the frozen fav2 voice (anchor + continuation, 10 diffusion steps), every clip gated by Whisper (text match), UTMOS ≥ 3.2 and speaker similarity ≥ 0.80, up to 3 takes per sentence. Source texts reviewed and minimally completed by Claude (cut sentences finished, numbers written as words).
- Matcha fine-tune: 10 k steps, batch 64, lr 5e-5, bf16, RTX 6000 Ada (≈ 1.5 h); quality plateaued from 5 k steps.
- Vocoder: teacher-forced mels of the 10 k checkpoint for all training clips paired with the VoxCPM2 audio, HiFi-GAN universal v1 fine-tuned 30 k steps, lr 2e-5.
Limitations
- One voice. The rhythm is more regular than the teacher's: Matcha's deterministic duration predictor averages timing.
- English detection uses a curated list (≈ 1 900 words and brands, Polish lookalikes excluded); an unlisted English word is read the Polish way unless you put it in braces.
- Synthetic voice: label generated speech as synthetic; don't present it as a real person.
Licence
CC BY-SA 4.0 (training data includes CC BY-SA texts; the base model is CC BY-SA 4.0). VoxCPM2 is Apache-2.0;
espeak-ng (runtime phonemizer) is GPL-3.0. See ATTRIBUTION.md.
Model tree for machinekind/Matcha-TTS-PL-v2
Base model
machinekind/Matcha-TTS-PL