phoonnx β€” Indic Parler-TTS (ONNX)

ONNX export of ai4bharat/indic-parler-tts for the phoonnx indic_parler engine.

Indic Parler-TTS speaks 20 Indic languages and English. You select the voice with a natural-language description ("Rohit's voice is clear and expressive..."), not with a speaker id or a reference clip.

Graphs

File Size Purpose
text_encoder.onnx 1.4 GB Flan-T5 encoder over the voice description
decoder_prefill.onnx 2.1 GB First AR step; emits self-attention and cross-attention KV
decoder_decode.onnx 1.3 GB Later AR steps; consumes both caches, emits self-attention KV only
dac_decoder.onnx 217 MB DAC 44.1 kHz codec decoder (9 codebooks)
tokenizer.json prompt tokenizer (the text you want spoken)
description_tokenizer.json description tokenizer (Flan-T5 vocabulary)
config.json self-describing phoonnx config (engine: indic_parler)

The two tokenizers are different vocabularies. Upstream is explicit about this: one tokenizer for the prompt, one for the description.

Architecture

description --> Flan-T5 encoder --> encoder states --> cross-attention (all 24 layers)
prompt      --> embed_prompts   --> prepended to the decoder input embeddings
decoder     --> 9 delayed DAC codebooks --> DAC decoder --> 44.1 kHz mono

Cross-attention keys and values do not change while decoding, so decoder_prefill computes them once and decoder_decode reads them back unchanged.

Parity

Every graph is float32 and was checked against the PyTorch model (parler_tts.ParlerTTSForConditionalGeneration, eager attention) on ser9 CPU:

Check Result
Text encoder, max abs diff (hi / en / ta) 3.9e-07 / 6.5e-06 / 3.2e-07
Prefill KV, max abs diff (self K/V, cross K/V) 8.1e-06 / 4.5e-06 / 2.4e-06 / 1.2e-06
Greedy codes, 60 steps identical, all 3 languages
Waveform correlation vs torch 0.9999999917 / 1.0000000000 / 1.0000000000
Waveform RMSE vs torch 1.3e-08 / 5.0e-08 / 1.5e-08

No quantised variants are published: the quantised graphs were not verified, and phoonnx does not ship unverified exports.

Notice

Upstream model: ai4bharat/indic-parler-tts by AI4Bharat, built on huggingface/parler-tts by Yoach Lacombe, Vaibhav Srivastav and Sanchit Gandhi. Apache-2.0, preserved from upstream. The DAC codec is Descript's, as vendored by upstream.

The weights are unmodified: the export wraps the PyTorch modules and traces them. No surgery, no quantisation.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for OpenVoiceOS/phoonnx-indic-parler

Quantized
(2)
this model

Collection including OpenVoiceOS/phoonnx-indic-parler