ByT5 multilingual G2P (tiny) β€” 17M params

Byte-level seq2seq, 142 language/ variety tags (eng-US/eng-UK, por-BR/por-PT, spa-ES/spa-LatAm, Welsh N/S, Armenian E/W, Bengali varieties, 20+ Sinitic splits). Trained on a 4.12M-pair harmonized corpus, with IPA conventions unified per language via a learned M2M aligner.

Input: <lang>: word (ISO-639-3, variety-suffixed where split β€” <eng-US>: hello, <spa-ES>: abeja). Output: space-separated IPA in gruut-anchored convention (h Ι› l ˈoʊ) β€” what Piper-family voices consume directly.

Results (4k stratified test sample, same test set)

this model previous (0.729 on old test)
micro exact 0.718 0.582 (same test, same eval)
macro exact 0.619 0.644

The tiny (17M) stays within striking distance of the small (300M) at 1/18th the size. The new corpus (+27% more entries, harmonized conventions) gives the tiny model a +13.6 pt lift on the same test set (0.582 β†’ 0.718).

Files

  • HF-format weights at root (~70 MB)
  • onnx/ β€” validated encoder+decoder pair (manual TorchScript export, 3/3 gate-pair validation passing). Consume with onnx_reference.py (a minimal correct consumer). CRITICAL conventions: token id = byte + 3; EOS appended to encoder input; decoder needs an explicit causal mask and a length-2 bootstrap β€” the reference script encodes all of them.

Quickstart (ONNX, CPU)

# pip install onnxruntime numpy
from huggingface_hub import snapshot_download
import sys; sys.path.insert(0, "voicegarden-lexicons/scripts/train_byt5")
from onnx_reference import load, run

d = snapshot_download("willwade/byt5-g2p-multilingual-tiny", allow_patterns=["onnx/*"])
enc, dec = load(d + "/onnx")
print(run(enc, dec, "<eng-US>: floravox"))

Training

The checkpoint was retrained on a 4.12M-pair harmonized corpus built from gruut + WikiPron sources with M2M-convention alignment. Training details in the voicegarden-lexicons repository (scripts/train_byt5/). Run: RTX 3090, ~8h, $1.20.

Licence

CC BY-SA 4.0 (share-alike inherited from WikiPron training data). Attribution: Wiktionary/WikiPron (CUNY-CL), gruut (rhasspy), Google byt5 base (Apache-2.0). Training code: voicegarden-lexicons/scripts/train_byt5.

Downloads last month
1,495
Safetensors
Model size
17.9M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for willwade/byt5-g2p-multilingual-tiny

Quantized
(6)
this model