loudkit / README.md
jer3mi's picture
Space: readable algorithm ID captions; README snippet on the 0.1.1 API
69a9113 verified
|
Raw History Blame Contribute Delete
4.1 kB
---
title: loudkit
emoji: ๐Ÿ”Š
colorFrom: gray
colorTo: red
sdk: gradio
sdk_version: 5.50.0
python_version: "3.12.12"
app_file: app.py
pinned: false
license: apache-2.0
short_description: On-device TTS. 28 voices, ten languages, two models.
models:
- loudreader/loudr-1
- loudreader/loudr-1-turbo
preload_from_hub:
- loudreader/loudr-1 loudr-1.safetensors,loudr-1-enrollment.safetensors,ve.safetensors,tokenizer.json,manifest.json,release.json,SHA256SUMS,voices/carmen.safetensors,voices/clara.safetensors,voices/colette.safetensors,voices/dante.safetensors,voices/darkman.safetensors,voices/dave.safetensors,voices/emma.safetensors,voices/freja.safetensors,voices/gosia.safetensors,voices/henri.safetensors,voices/henry.safetensors,voices/ines.safetensors,voices/joe.safetensors,voices/kathleen.safetensors,voices/kerstin.safetensors,voices/lucy.safetensors,voices/miles.safetensors,voices/nathalie.safetensors,voices/nils.safetensors,voices/oliver.safetensors,voices/oscar.safetensors,voices/paola.safetensors,voices/pim.safetensors,voices/selma.safetensors,voices/sophie.safetensors,voices/soren.safetensors,voices/thorsten.safetensors,voices/tugao.safetensors 7516673ee74ab2228ffc4e99223d296d8904d7ff
- loudreader/loudr-1-turbo loudr-1-turbo.safetensors,tokenizer.json,manifest.json,release.json,SHA256SUMS,voices/carmen.safetensors,voices/clara.safetensors,voices/colette.safetensors,voices/dante.safetensors,voices/darkman.safetensors,voices/dave.safetensors,voices/emma.safetensors,voices/freja.safetensors,voices/gosia.safetensors,voices/henri.safetensors,voices/henry.safetensors,voices/ines.safetensors,voices/joe.safetensors,voices/kathleen.safetensors,voices/kerstin.safetensors,voices/lucy.safetensors,voices/miles.safetensors,voices/nathalie.safetensors,voices/nils.safetensors,voices/oliver.safetensors,voices/oscar.safetensors,voices/paola.safetensors,voices/pim.safetensors,voices/selma.safetensors,voices/sophie.safetensors,voices/soren.safetensors,voices/thorsten.safetensors,voices/tugao.safetensors 366f140a5f7ebe215e5b43cdd73384746c83bae6
---
# loudkit
Twenty-eight voices across ten languages, on two models:
[loudreader/loudr-1](https://huggingface.co/loudreader/loudr-1) and
[loudreader/loudr-1-turbo](https://huggingface.co/loudreader/loudr-1-turbo),
running the [loudkit](https://github.com/loudreader/loudkit) engine on ZeroGPU.
## Two tabs
- **Voices.** Twenty-eight voices, each beside the reference recording it was
enrolled from, and each rendered on both models from the same passage at the
same seed. English first, with a dropdown for the other nine languages. The
samples were rendered ahead of time and ship in this repo, so playing them
uses no GPU. Under them is a box for your own text, up to 1,000 characters,
and one that renders the same words on both models side by side.
- **Clone.** A voice from about ten seconds of a recording, then speak with it
on either model. One profile serves both, which is what a portable voice
profile means. The microphone is the default path; an upload is gated on a
confirmation, and neither recording outlives the request.
## What this demo is honest about
ZeroGPU bills GPU time to the visitor. Listening costs nothing, because the
fifty-six samples are files. Speaking and cloning spend the visitor's own daily
quota, which is why the text boxes are capped well below what the library takes.
Both models are pinned to an immutable Hub commit rather than to a branch, so
the algorithm ID the page prints belongs to the bytes it ran.
Output files carry
[unsigned loudkit provenance metadata](https://github.com/loudreader/loudkit/blob/main/docs/reference/provenance.md):
the algorithm ID, the recipe and the seed.
## Running it locally instead
```bash
pip install "loudkit[torch,audio,enroll,hub]"
```
```python
import loudkit as lk
engine = lk.load("loudreader/loudr-1", revision="v0.1.1")
voice = lk.voice("kathleen", repo="loudreader/loudr-1", revision="v0.1.1")
engine.synthesize("Hello from loudkit.", voice, seed=7).save("hello.wav")
```
Nothing is queued and nothing is metered.