KittenTTS 2

A 1.7B speech language model with in-context voice cloning. It reads text and writes S3 codec tokens, which a vocoder turns into 24 kHz audio.

pip install kittenml
from kittenml import KittenTTS
import soundfile as sf

m = KittenTTS("KittenML/kitten-tts-2")

# a built-in voice
audio = m.generate("One day, a little girl named Lily found a needle in her room.",
                   voice="Bruno")

# or clone one, from 5-30 seconds of a single speaker
audio = m.generate("This is my own voice.", reference="my_voice.wav")

sf.write("output.wav", audio, m.sample_rate)

Everything the model needs is in this repository, so no Hugging Face login is required.

Voices

m.available_voices lists all 47. Bella, Jasper, Luna, Bruno, Rosie, Hugo, Kiki and Leo are the same speakers as in KittenTTS 0.8, so existing code keeps working.

Nine are named after a language rather than a person — Arabic, Chinese, French, German, Hindi, Italian, Portuguese, Russian, Spanish — and are how you reach those languages, since the voice is what carries the accent:

m.generate("Guten Morgen. Ich wünsche dir einen wunderschönen Tag.",
           voice="German", normalize=False)

Pass normalize=False for non-English text: the text normalizer is English-tuned and will mangle numbers and dates in other languages.

Expression controls

Beta. Emotion control steers delivery rather than guaranteeing it, and the effect varies by voice and by sentence.

m.generate("[joyful] We won the grant <laugh> I can (((hardly))) believe it!",
           voice="Kiki", preset="expressive")

A leading [emotion] tag, inline <event> tags and (((emphasis))) spans reach the model as markup rather than being spoken, and switch on its expression conditioning.

Emotions — one leading tag sets the emotion for the whole line:

[angry] [contemplative] [excited] [joyful] [mundane] [nervous] [sad] [stern] [surprised] [tender]

Vocal events — inline, anywhere in the line:

<gasp> <giggle> <growl> <gulp> <laugh> <pause> <scoff> <sigh> <sob> <um>

Emphasis — triple parentheses stress a word or short phrase:

m.generate("I told you (((never))) to open that door.", voice="Victor")

Only these twenty tags are recognised. They are the most common of the many in the training data, so a rarer one such as [reverent] is spoken as ordinary text rather than treated as markup — as is anything else bracketed, like section [3] or x < 5.

Weight variants

The language model ships in three packings of the same weights. weights= picks one; model.available_weights lists them.

weights= On disk
"packed" (default) 954 MB 1.58-bit ternary body, bf16 embedding. Lossless
"emb4" 469 MB Same body in TL2, plus a 4-bit embedding. Lossy
"full" 3469 MB Plain bf16
m = KittenTTS("KittenML/kitten-tts-2", weights="emb4")

Only the variant you ask for is downloaded.

emb4 halves the download, and the saving is almost entirely the token embedding — 324M parameters that packed has to leave at bf16 because they are not ternary. Its transformer body is still bit-exact; the embedding is not, at L2 relative error 0.118 against bf16. Measured at export, that costs roughly a tenth of a point of perplexity on internal evaluations. Small, but it is the one lossy thing here, which is why packed stays the default.

Decoders

Audio is decoded in two stages, and the first can be swapped for a smaller distilled student with weights packed to 4 or 8 bits:

m = KittenTTS("KittenML/kitten-tts-2", decoder="student_w4")
Decoder Flow on disk
default 459 MB Best quality
student_w4 39 MB Distilled single-step student, weights packed to 4 bits

These trade fidelity for footprint, not for speed: quantisation shrinks storage and memory bandwidth, not arithmetic.

What is in here

lm/          the speech language model, plus its spk_proj speaker head
speaker/     the speaker-embedding model used when cloning
voices/      reference clips, transcripts, and precomputed embeddings
decoders/    the optional 4-bit decoder
cpp/         GGUF weights for the llama.cpp fork, see "Running on CPU"
config.json  token layout, decode presets, voice and decoder indexes

The weights are 947 MiB: 910 MiB for the language model and 37 MiB for the decoder. A load pulls those rather than the whole repository, and only the decoder you ask for.

The language model's linear weights are ternary — within every 128-wide group each value is exactly one of {-scale, 0, +scale} — so bf16 spends 16 bits to say one of three things. lm/model-ternary.safetensors packs them five trits to a byte, 1.6 bits per weight, with each group keeping its scale at full precision:

lm/model.safetensors 3.47 GB, bf16 throughout
lm/model-ternary.safetensors 0.95 GB, the same weights, 3.6x smaller

The packing is exact rather than approximate, so the two produce bit-identical audio. config.json points lm_packed at the smaller file and that is what gets downloaded; the full file stays for anything loading this with plain transformers.

Running on CPU

KittenTTS 2 runs on CPU out of the box — device is auto-detected — but the fastest way is kitten-tts-2-cpp, our llama.cpp fork. The cpp/ directory here holds what it needs: the GGUF language model and its decoders.

Requirements

Python 3.10 or later, and PyTorch.

License

These weights are released under the Stellon Labs Community License. Research, non-commercial and limited commercial use are free of charge; the commercial grant ends once you or your affiliates pass USD $1,000,000 in annual revenue or in total cumulative funding, at which point you need a separate license from Stellon Labs. Distributing the weights, a derivative, or a product built on them carries attribution requirements — see Section IV(a).

Two things in here are covered by their own terms instead: speaker/, the pyannote speaker-embedding model, under MIT, and the kittenml Python package that loads this repository, under Apache 2.0.

Downloads last month
55
GGUF
Model size
2B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support