loudkit / README.md
jer3mi's picture
Space: readable algorithm ID captions; README snippet on the 0.1.1 API
69a9113 verified
|
Raw History Blame Contribute Delete
4.1 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: loudkit
emoji: 🔊
colorFrom: gray
colorTo: red
sdk: gradio
sdk_version: 5.50.0
python_version: 3.12.12
app_file: app.py
pinned: false
license: apache-2.0
short_description: On-device TTS. 28 voices, ten languages, two models.
models:
  - loudreader/loudr-1
  - loudreader/loudr-1-turbo
preload_from_hub:
  - >-
    loudreader/loudr-1
    loudr-1.safetensors,loudr-1-enrollment.safetensors,ve.safetensors,tokenizer.json,manifest.json,release.json,SHA256SUMS,voices/carmen.safetensors,voices/clara.safetensors,voices/colette.safetensors,voices/dante.safetensors,voices/darkman.safetensors,voices/dave.safetensors,voices/emma.safetensors,voices/freja.safetensors,voices/gosia.safetensors,voices/henri.safetensors,voices/henry.safetensors,voices/ines.safetensors,voices/joe.safetensors,voices/kathleen.safetensors,voices/kerstin.safetensors,voices/lucy.safetensors,voices/miles.safetensors,voices/nathalie.safetensors,voices/nils.safetensors,voices/oliver.safetensors,voices/oscar.safetensors,voices/paola.safetensors,voices/pim.safetensors,voices/selma.safetensors,voices/sophie.safetensors,voices/soren.safetensors,voices/thorsten.safetensors,voices/tugao.safetensors
    7516673ee74ab2228ffc4e99223d296d8904d7ff
  - >-
    loudreader/loudr-1-turbo
    loudr-1-turbo.safetensors,tokenizer.json,manifest.json,release.json,SHA256SUMS,voices/carmen.safetensors,voices/clara.safetensors,voices/colette.safetensors,voices/dante.safetensors,voices/darkman.safetensors,voices/dave.safetensors,voices/emma.safetensors,voices/freja.safetensors,voices/gosia.safetensors,voices/henri.safetensors,voices/henry.safetensors,voices/ines.safetensors,voices/joe.safetensors,voices/kathleen.safetensors,voices/kerstin.safetensors,voices/lucy.safetensors,voices/miles.safetensors,voices/nathalie.safetensors,voices/nils.safetensors,voices/oliver.safetensors,voices/oscar.safetensors,voices/paola.safetensors,voices/pim.safetensors,voices/selma.safetensors,voices/sophie.safetensors,voices/soren.safetensors,voices/thorsten.safetensors,voices/tugao.safetensors
    366f140a5f7ebe215e5b43cdd73384746c83bae6

loudkit

Twenty-eight voices across ten languages, on two models: loudreader/loudr-1 and loudreader/loudr-1-turbo, running the loudkit engine on ZeroGPU.

Two tabs

  • Voices. Twenty-eight voices, each beside the reference recording it was enrolled from, and each rendered on both models from the same passage at the same seed. English first, with a dropdown for the other nine languages. The samples were rendered ahead of time and ship in this repo, so playing them uses no GPU. Under them is a box for your own text, up to 1,000 characters, and one that renders the same words on both models side by side.
  • Clone. A voice from about ten seconds of a recording, then speak with it on either model. One profile serves both, which is what a portable voice profile means. The microphone is the default path; an upload is gated on a confirmation, and neither recording outlives the request.

What this demo is honest about

ZeroGPU bills GPU time to the visitor. Listening costs nothing, because the fifty-six samples are files. Speaking and cloning spend the visitor's own daily quota, which is why the text boxes are capped well below what the library takes.

Both models are pinned to an immutable Hub commit rather than to a branch, so the algorithm ID the page prints belongs to the bytes it ran.

Output files carry unsigned loudkit provenance metadata: the algorithm ID, the recipe and the seed.

Running it locally instead

pip install "loudkit[torch,audio,enroll,hub]"
import loudkit as lk

engine = lk.load("loudreader/loudr-1", revision="v0.1.1")
voice = lk.voice("kathleen", repo="loudreader/loudr-1", revision="v0.1.1")
engine.synthesize("Hello from loudkit.", voice, seed=7).save("hello.wav")

Nothing is queued and nothing is metered.