mbfa-latin (p4mbfa · latin)
A context-free CNN phoneme encoder and forced aligner (p4mbfa, the successor of CUPE / p3cupe), for the latin language group of standard_g2p. Each 5 ms frame is classified from at most 120 ms of audio (38.9 ms receptive field), so the model cannot learn any language's phonotactics. Phone boundaries come from a segmental Viterbi decoder over the known phone sequence, refined to sub-frame precision at the crossing of neighbouring posteriors.
| head | labels | size |
|---|---|---|
ph |
local tokens of latin (incl. <blank> SIL noise <unk>) |
208 |
phg |
gold phoneme groups, shared by every language group | 15 |
tone |
tone values, shared by every language group; present but untrained (no language in this group has a tone layer) | 22 |
Languages (the FLEURS languages it was trained on): af, az, bs, ca, cs, cy, da, de, en-US, es-419, et, fi, fil, fr, ga, gl, hu, id, is, it, lb, lt, mi, ms, mt, nb, nl, pl, pt-BR, ro, sk, sl, sv, sw, tr, uz. standard_g2p maps every member of the group onto the same tokens, so other members may align too, untested.
Training
Data: FLEURS latin, all 273.5 h, training noise_level 0.05 (~26 dB SNR).
Checkpoint: experiment ma04a, epoch 6, selected on TIMIT f1@25ms on the tune half of the 400-clip probe, best of the noise-0.05 latin runs.
The recommended latin model: trained with the most noise of any latin run (0.05), for real recordings. Continued from ma02a's final epoch (ramp_frames 1.0). TIMIT is clean speech, so it scores lower there than mbfa-latin-benchmark (72.93 holdout), which trained at noise 0.02.
| metric | value |
|---|---|
timit_f1_25ms_tune |
71.22 |
timit_f1_25ms_holdout |
71.7 |
timit_f1_25ms_nobonus_holdout |
65.21 |
timit_f1_10ms_holdout |
53.99 |
timit_abs_err_ms_holdout |
16.13 |
timit_*: 400 TIMIT clips (English, never trained on). Each clip's text is phonemised with standard_g2p, force-aligned, and the predicted onsets are scored against TIMIT's hand-labelled phone boundaries: F1 within 25 / 10 ms, and the mean absolute error of matched boundaries. The clips are split in two halves: tune picked the checkpoint, holdout did not. nobonus aligns without the log-mel onset bonus, which shows the network's own share.
Labels are standard_g2p dictionary pronunciations (gold inventory 9438371ed6dd),
not phonetic transcriptions of what was said.
Usage
The code is in https://github.com/tabahi/bfa_models (p4mbfa/); clone it and run from its root.
from p4mbfa.inference import MbfaAligner
aligner = MbfaAligner.from_pretrained("Tabahi/mbfa-latin")
wav = aligner.load_audio("clip.wav") # mono, 16000 Hz
phones = ["SIL", "h", "ɛ", "l", "o", "SIL"] # this group's tokens (config.json labels.tokens)
for seg in aligner.align(wav, phones):
print(seg["token"], seg["start_ms"], seg["end_ms"])
aligner.encode(wav) returns the per-frame log posteriors of all three heads.
Text -> tokens is standard_g2p's job (goldG2P.phonemize_sentence then
lang_group_inventory.to_local); config.json lists the token strings. from_pretrained
downloads only config.json and model.safetensors.
Fine-tuning
ckpt/latin_fleurs273h_ma04a_e6_timit_f1_25ms=71.22.ckpt is the training checkpoint (PyTorch Lightning, weights only, a pickle: load
it only if you trust this repo). Download it, then point a p4mbfa yaml at it to continue
training, or to train a new language group on its trunk and shared phg head:
hf download Tabahi/mbfa-latin "ckpt/latin_fleurs273h_ma04a_e6_timit_f1_25ms=71.22.ckpt" --local-dir tmp/hf/mbfa-latin
ckpt_path: "tmp/hf/mbfa-latin/ckpt/latin_fleurs273h_ma04a_e6_timit_f1_25ms=71.22.ckpt"
reset_fine_heads: true # new lang_group: rebuild ph_head / tone_head; false = same group
License: AGPL-3.0.
- Downloads last month
- -