mbfa-cjk (p4mbfa · cjk)

A context-free CNN phoneme encoder and forced aligner (p4mbfa, the successor of CUPE / p3cupe), for the cjk language group of standard_g2p. Each 5 ms frame is classified from at most 120 ms of audio (38.9 ms receptive field), so the model cannot learn any language's phonotactics. Phone boundaries come from a segmental Viterbi decoder over the known phone sequence, refined to sub-frame precision at the crossing of neighbouring posteriors.

head labels size
ph local tokens of cjk (incl. <blank> SIL noise <unk>) 48
phg gold phoneme groups, shared by every language group 15
tone tone values, shared by every language group; trained on this group's tone layer 22

Languages (the FLEURS languages it was trained on): cmn, yue. standard_g2p maps every member of the group onto the same tokens, so other members may align too, untested.

Training

Data: FLEURS cjk, a 62% sample (train_limit 100000): ~9 of 14.9 h, training noise_level 0.02.
Checkpoint: experiment mc01a, epoch 13, selected on FLEURS val_loss (best of 30 epochs).
Trunk and phg_head from latin ma02a (final epoch), ph_head and tone_head trained fresh.

metric value
val_loss 2.5963
val_frame_acc 0.4795
val_frame_acc_groups 0.612

val_* are on held-out FLEURS clips but are scored against the model's own alignments (there are no boundary labels outside English), so they measure self-consistency, not boundary accuracy.

Labels are standard_g2p dictionary pronunciations (gold inventory 9438371ed6dd), not phonetic transcriptions of what was said.

Usage

The code is in https://github.com/tabahi/bfa_models (p4mbfa/); clone it and run from its root.

from p4mbfa.inference import MbfaAligner

aligner = MbfaAligner.from_pretrained("Tabahi/mbfa-cjk")
wav = aligner.load_audio("clip.wav")                     # mono, 16000 Hz
phones = ["SIL", "h", "ɛ", "l", "o", "SIL"]   # this group's tokens (config.json labels.tokens)
for seg in aligner.align(wav, phones):
    print(seg["token"], seg["start_ms"], seg["end_ms"])

aligner.encode(wav) returns the per-frame log posteriors of all three heads. Text -> tokens is standard_g2p's job (goldG2P.phonemize_sentence then lang_group_inventory.to_local); config.json lists the token strings. from_pretrained downloads only config.json and model.safetensors.

Fine-tuning

ckpt/cjk_fleurs9h_mc01a_e13_val_loss=2.596.ckpt is the training checkpoint (PyTorch Lightning, weights only, a pickle: load it only if you trust this repo). Download it, then point a p4mbfa yaml at it to continue training, or to train a new language group on its trunk and shared phg head:

hf download Tabahi/mbfa-cjk "ckpt/cjk_fleurs9h_mc01a_e13_val_loss=2.596.ckpt" --local-dir tmp/hf/mbfa-cjk
ckpt_path: "tmp/hf/mbfa-cjk/ckpt/cjk_fleurs9h_mc01a_e13_val_loss=2.596.ckpt"
reset_fine_heads: true      # new lang_group: rebuild ph_head / tone_head; false = same group

License: AGPL-3.0.

Downloads last month
-
Safetensors
Model size
14.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Tabahi/mbfa-cjk