--- title: loudkit emoji: 🔊 colorFrom: gray colorTo: red sdk: gradio sdk_version: 5.50.0 python_version: "3.12.12" app_file: app.py pinned: false license: apache-2.0 short_description: On-device TTS. 28 voices, ten languages, two models. models: - loudreader/loudr-1 - loudreader/loudr-1-turbo preload_from_hub: - loudreader/loudr-1 loudr-1.safetensors,loudr-1-enrollment.safetensors,ve.safetensors,tokenizer.json,manifest.json,release.json,SHA256SUMS,voices/carmen.safetensors,voices/clara.safetensors,voices/colette.safetensors,voices/dante.safetensors,voices/darkman.safetensors,voices/dave.safetensors,voices/emma.safetensors,voices/freja.safetensors,voices/gosia.safetensors,voices/henri.safetensors,voices/henry.safetensors,voices/ines.safetensors,voices/joe.safetensors,voices/kathleen.safetensors,voices/kerstin.safetensors,voices/lucy.safetensors,voices/miles.safetensors,voices/nathalie.safetensors,voices/nils.safetensors,voices/oliver.safetensors,voices/oscar.safetensors,voices/paola.safetensors,voices/pim.safetensors,voices/selma.safetensors,voices/sophie.safetensors,voices/soren.safetensors,voices/thorsten.safetensors,voices/tugao.safetensors 7516673ee74ab2228ffc4e99223d296d8904d7ff - loudreader/loudr-1-turbo loudr-1-turbo.safetensors,tokenizer.json,manifest.json,release.json,SHA256SUMS,voices/carmen.safetensors,voices/clara.safetensors,voices/colette.safetensors,voices/dante.safetensors,voices/darkman.safetensors,voices/dave.safetensors,voices/emma.safetensors,voices/freja.safetensors,voices/gosia.safetensors,voices/henri.safetensors,voices/henry.safetensors,voices/ines.safetensors,voices/joe.safetensors,voices/kathleen.safetensors,voices/kerstin.safetensors,voices/lucy.safetensors,voices/miles.safetensors,voices/nathalie.safetensors,voices/nils.safetensors,voices/oliver.safetensors,voices/oscar.safetensors,voices/paola.safetensors,voices/pim.safetensors,voices/selma.safetensors,voices/sophie.safetensors,voices/soren.safetensors,voices/thorsten.safetensors,voices/tugao.safetensors 366f140a5f7ebe215e5b43cdd73384746c83bae6 --- # loudkit Twenty-eight voices across ten languages, on two models: [loudreader/loudr-1](https://huggingface.co/loudreader/loudr-1) and [loudreader/loudr-1-turbo](https://huggingface.co/loudreader/loudr-1-turbo), running the [loudkit](https://github.com/loudreader/loudkit) engine on ZeroGPU. ## Two tabs - **Voices.** Twenty-eight voices, each beside the reference recording it was enrolled from, and each rendered on both models from the same passage at the same seed. English first, with a dropdown for the other nine languages. The samples were rendered ahead of time and ship in this repo, so playing them uses no GPU. Under them is a box for your own text, up to 1,000 characters, and one that renders the same words on both models side by side. - **Clone.** A voice from about ten seconds of a recording, then speak with it on either model. One profile serves both, which is what a portable voice profile means. The microphone is the default path; an upload is gated on a confirmation, and neither recording outlives the request. ## What this demo is honest about ZeroGPU bills GPU time to the visitor. Listening costs nothing, because the fifty-six samples are files. Speaking and cloning spend the visitor's own daily quota, which is why the text boxes are capped well below what the library takes. Both models are pinned to an immutable Hub commit rather than to a branch, so the algorithm ID the page prints belongs to the bytes it ran. Output files carry [unsigned loudkit provenance metadata](https://github.com/loudreader/loudkit/blob/main/docs/reference/provenance.md): the algorithm ID, the recipe and the seed. ## Running it locally instead ```bash pip install "loudkit[torch,audio,enroll,hub]" ``` ```python import loudkit as lk engine = lk.load("loudreader/loudr-1", revision="v0.1.1") voice = lk.voice("kathleen", repo="loudreader/loudr-1", revision="v0.1.1") engine.synthesize("Hello from loudkit.", voice, seed=7).save("hello.wav") ``` Nothing is queued and nothing is metered.