Paradee-8M v1.0

Paradee is a small English text-to-speech model. It has 8.07M parameters and is distilled from Kokoro-82M, and it speaks one voice, Kokoro's af_heart.

  • Small. The int8 model is one 9 MB ONNX file, against 325 MB for Kokoro.
  • Fast. It runs about 18x faster than real time on one CPU thread, with no GPU.
  • Close to its teacher. It scores 4.41 on UTMOS (Kokoro: 4.52) and has the same word error rate (5.7%).

Paper: Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Code, training and the Python package: github.com/sahilmahendrakar/paradee

Samples

Five held-out sentences, read by Paradee and by its teacher. The sentences are in samples/sentences.txt.

# Paradee Kokoro-82M (teacher)
1 listen listen
2 listen listen
3 listen listen
4 listen listen
5 listen listen

Usage

pip install git+https://github.com/sahilmahendrakar/paradee
python -m paradee "Paradee is a small voice that runs anywhere." -o hello.wav
from paradee import Paradee, SAMPLE_RATE
import soundfile as sf

tts = Paradee()
sf.write("hello.wav", tts("Paradee is a small voice that runs anywhere."), SAMPLE_RATE)

Files

File What it is
onnx/paradee_int8.onnx The whole model in one graph, weights in int8 (9.0 MB). Use this one.
onnx/paradee.onnx The same graph in fp32 (37 MB). It sounds the same.
config.json Kokoro's phoneme vocabulary and the sample rate
tokenizer.json The same vocabulary in the format transformers.js and kokoro-js load
pytorch/text_side.pt, pytorch/decoder.pt PyTorch weights for the two halves, for the training code on GitHub
samples/ The audio above

The ONNX graph goes from phoneme ids to audio. It includes the phase-locking filter described below.

  • Inputs: input_ids, int64 [1, T], phoneme ids from config.json with a 0 pad token at each end, at most 512 in total. The second input is speed, float32 [1], where 1.0 is normal speed.
  • Output: waveform, float32 [1, samples] at 24 kHz.

Phonemes must be written the way misaki, Kokoro's own grapheme-to-phoneme library, writes them, because that is all Paradee saw in training. In the browser, kokoro-js phonemizes with eSpeak NG instead. There, web/misaki.js converts eSpeak's spelling to misaki's. Without it, Whisper mishears about 31% of words, and with it about 2%.

Results

All numbers are on 200 held-out sentences. Speed is on one CPU thread of an Apple M4 Pro.

Model (voice) Params File size Speed UTMOS WER
Kokoro-82M, teacher (af_heart) 81.8M 325 MB 7.6x 4.52 5.7%
Paradee (af_heart) 8.07M 8.45 MB 25.0x 4.41 5.7%
Kokoro-7M-Distill (af_msa) 7.48M 30.1 MB 35.5x 4.18 7.4%
Piper, en_US-lessac-medium 15.7M 63.2 MB 15.4x 4.36 8.8%
KittenTTS nano 0.8 (Bella) 14.0M 56.8 MB 10.7x 4.01 5.5%

UTMOS is a neural network trained on human ratings. It predicts how natural a clip sounds, on a scale from 1 to 5. WER is the share of words that Whisper (base) transcribes wrongly.

The table is from the paper and measures PyTorch. The released paradee_int8.onnx scores UTMOS 4.41 on the same sentences and runs about 18x faster than real time in onnxruntime on one thread.

How it was made

Paradee is Kokoro's own code at smaller widths: a 4.23M text side and a 3.85M decoder. The two halves were trained separately against the frozen teacher, on 12,000 WikiText-103 sentences (23.9 hours) read by Kokoro.

  1. The text side learns to predict the teacher's phoneme durations, pitch, loudness and phoneme features.
  2. The decoder learns to turn the teacher's saved values into the teacher's audio, first with spectrogram losses and then with adversarial training.
  3. The halves are joined with no further training.
  4. A filter with no parameters corrects the phase of voiced sound between 2 and 8 kHz. Phase is the timing of each frequency's wave. This removes a slight buzz that the small decoder otherwise leaves.

All training ran on one MacBook Pro.

Limitations

  • English only, with American pronunciation, and one voice.
  • Numbers, abbreviations and unusual words are pronounced only as well as misaki handles them.
  • A faint buzz can still be heard on some voiced sounds, though much less than without the filter.

License

Apache 2.0, the same as Kokoro-82M.

Citation

@misc{mahendrakar2026paradee,
  title  = {Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model},
  author = {Mahendrakar, Sahil},
  year   = {2026},
  url    = {https://huggingface.co/sahilmahendrakar/Paradee-8M-v1.0}
}
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TechnoBaptist/Paradee-8M-v1.0

Quantized
(85)
this model

Paper for TechnoBaptist/Paradee-8M-v1.0