Matcha-TTS-PL v2 — one synthetic voice, modern Polish, English words

Second version of Matcha-TTS-PL: the same small, fast, non-autoregressive Matcha-TTS (20.9 M parameters, real-time on a small GPU or in a browser) distilled from VoxCPM2, a 2B-parameter TTS, onto a single synthetic voice. What changes compared with v1:

  • One original voice, fav2 — designed from a text description with VoxCPM2 (OpenBMB, Apache-2.0) and frozen as a prompt cache; not a recording or a clone of a real person (closest of our reference readers: speaker similarity 0.83).
  • Modern Polish of a humanoid assistant — 13.1 h of training speech: assistant speech in 18 everyday domains, modern corpora (NKJP, Wikinews, KPWr, ParlaMint, OpenAssistant, Wikibooks, Wikivoyage) and reviewed YouTube transcripts; two thirds of it in multi-sentence passages (10–25 s), so prosody is planned across sentences.
  • English words inside Polish sentences — "Po meetingu wyślę ci feedback na maila" is read with English phonemes: English words are detected automatically and phonemized by espeak-ng en-us, the rest by espeak-ng pl (english/). No pronunciation dictionary needed for common loanwords and brands.
  • A HiFi-GAN fine-tuned on this voice, with real ground truth (the teacher's audio) — the v1 vocoder was tuned for the audiobook readers and blurs this voice.

Try it in the browser: spaces/machinekind/matcha-tts-pl-v2.

Quality (12 held-out sentences, Mac CPU, 4 ODE steps, temperature 0.8)

system UTMOS ↑ Whisper WER ↓ similarity to fav2 ↑
VoxCPM2 teacher (2B, autoregressive) 3.68 1.3 % 0.883
v2 (this model + its vocoder) 3.40 1.3 % 0.850
v1 model with the new voice slot, before fine-tuning 2.91 2.7 % 0.725

WER counts Whisper's spelling too: the teacher's two "errors" are "krutki" for krótki and "HWSR" for the robot's name. Speaking rate: the released default tempo is 0.8 (length_scale); at 1.0 v2 speaks about 19 % slower than the teacher.

Files

path what
onnx/matcha_pl_t4.onnx, onnx/matcha_pl_t2.onnx acoustic model + fav2 vocoder in one graph, 4 / 2 ODE steps; inputs x (phoneme ids), x_lengths, scales=[temperature, length_scale], spk_emb (float32 [1, 64]); outputs wav (22.05 kHz), wav_lengths
onnx/voices.json the fav2 speaker embedding, 18 style embeddings, defaults
model/matcha_pl_v2.ckpt PyTorch checkpoint (Matcha-TTS + the tts-pl patch: style tokens, Polish cleaners) — speaker 20 = fav2
vocoder/generator_fav2, vocoder/config.json HiFi-GAN (v1 architecture) fine-tuned on this model's mels ↔ the fav2 audio
english/ English detection (english_detect.py, english_words.txt; browser: english.js, english_data.json) and the mixed phonemizer (polish_mixed_cleaners.py; browser: cleaner.js)
samples/ the 12 held-out sentences, v2 and teacher
data/ style token map, speaker map
ATTRIBUTION.md every text source of the training data, and the base model's data

Quick start (ONNX, Python)

import sys, json, numpy as np, onnxruntime as ort, soundfile as sf
sys.path.append("english")                          # english_detect.py + english_words.txt (optional fallback: pip install wordfreq)
from english_detect import mark_english
from matcha.text import text_to_sequence            # Matcha-TTS with the tts-pl cleaners: paste english/polish_cleaners.py and
from matcha.utils.utils import intersperse          # english/polish_mixed_cleaners.py into matcha/text/cleaners.py
text = mark_english("Po meetingu wyślę ci feedback na maila, a w weekend zrobię update.")
ids = np.array(intersperse(text_to_sequence(text, ["polish_mixed_cleaners"])[0], 0), np.int64)
vj = json.load(open("onnx/voices.json"))
spk = np.array(vj["speakers"]["0"]["emb"], np.float32) + np.array(vj["styles"]["7"]["emb"], np.float32)   # 7 = neutral style
sess = ort.InferenceSession("onnx/matcha_pl_t4.onnx")
wav, n = sess.run(None, {"x": ids[None], "x_lengths": np.array([ids.size]), "scales": np.array([0.8, 0.8], np.float32), "spk_emb": spk[None]})
sf.write("out.wav", wav[0, :n[0]], 22050)   # scales = [temperature, length_scale (tempo)]

Text front end: mark_english() wraps English words in braces (Po {meeting}u wyślę ci {feedback}.), the cleaner phonemizes braced spans with espeak-ng en-us and the rest with pl (with the Polish ending of an inflected English word glued on unstressed, and one-letter prepositions before English words handled by rule). You can brace any word yourself. The browser port produces the training phonemes character for character (tested on 300 sentences).

Training

  • Base: Matcha-TTS-PL v1 (target model), new speaker row initialised from the closest real reader; same mel statistics, style tokens and phoneme table (all en-us phones were already in the symbol table, now trained).
  • Data: 6 493 clips, 13.1 h, generated by VoxCPM2 with the frozen fav2 voice (anchor + continuation, 10 diffusion steps), every clip gated by Whisper (text match), UTMOS ≥ 3.2 and speaker similarity ≥ 0.80, up to 3 takes per sentence. Source texts reviewed and minimally completed by Claude (cut sentences finished, numbers written as words).
  • Matcha fine-tune: 10 k steps, batch 64, lr 5e-5, bf16, RTX 6000 Ada (≈ 1.5 h); quality plateaued from 5 k steps.
  • Vocoder: teacher-forced mels of the 10 k checkpoint for all training clips paired with the VoxCPM2 audio, HiFi-GAN universal v1 fine-tuned 30 k steps, lr 2e-5.

Limitations

  • One voice. The rhythm is more regular than the teacher's: Matcha's deterministic duration predictor averages timing.
  • English detection uses a curated list (≈ 1 900 words and brands, Polish lookalikes excluded); an unlisted English word is read the Polish way unless you put it in braces.
  • Synthetic voice: label generated speech as synthetic; don't present it as a real person.

Licence

CC BY-SA 4.0 (training data includes CC BY-SA texts; the base model is CC BY-SA 4.0). VoxCPM2 is Apache-2.0; espeak-ng (runtime phonemizer) is GPL-3.0. See ATTRIBUTION.md.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for machinekind/Matcha-TTS-PL-v2

Quantized
(1)
this model

Space using machinekind/Matcha-TTS-PL-v2 1