CUPE v3 (p3cupe) 路 English
A context-free CNN phoneme encoder (the third CUPE), trained without boundary labels to place phoneme boundaries: one ph66 phoneme distribution every 5 ms, each frame seeing at most 120 ms of audio (38.9 ms receptive field). Boundaries come from a segmental Viterbi decoder over the known phone sequence, read at the crossing of neighbouring posteriors.
Superseded by p4mbfa (Tabahi/mbfa-<language group>, standard_g2p labels, many languages).
Kept for BFA's ph66 pipeline.
| head | labels | size |
|---|---|---|
| phonemes | ph66 (config.json labels.phonemes, index 0 = SIL; the last slot, 'noise', is never trained) |
67 |
| groups | ph66 phoneme groups (labels.phoneme_groups) | 17 |
Training
- Data: LibriSpeech, libri_500h (~497 h), one epoch.
- Checkpoint: experiment
pa07f, epoch 0, selected on TIMIT lev f1@25ms on the tune half, re-scored with the final decoder.
Warm-started from pa07e; the lineage is pa01d -> pa07a (MLS, 7 languages, 100 h) -> pa07b (libri_360h) -> pa07c, pa07d, pa07e (libri_500h) -> pa07f. Trained with min_phone_frames 3, soft 10, onset bonus 30, ramp 3. For per-phone onset accuracy decode with boundary_bonus_weight 0 (phone_hit 73.07 vs 68.73).
| metric | value |
|---|---|
timit_lev_f1_25ms_tune |
76.1 |
timit_lev_f1_25ms_holdout |
75.58 |
timit_lev_f1_25ms_bonus0_holdout |
73.77 |
timit_phone_hit_25ms_bonus0_holdout |
73.07 |
timit_phone_hit_25ms_holdout |
68.73 |
TIMIT: 400 clips (English, never trained on), split in a tune half (picked the checkpoint and
the decoder) and a holdout half. lev_f1_25ms is phone-boundary F1 within 25 ms;
phone_hit_25ms is the share of phones whose onset is within 25 ms. Scored with the decoder in
config.json ("decoder": min 6 + soft 10 frames,
log-mel onset bonus 30.0); bonus0 = the same decoder without the bonus.
The settings it was trained with are under training.decoder_at_training.
Usage
The code is in https://github.com/tabahi/bfa_models (p3cupe/); clone it and run from its root.
import torch
from p3cupe.model import CupeExtractor
from p3cupe.windowing import slice_windows, stitch_windows
ext = CupeExtractor("Tabahi/CUPE-3-en", device="cpu") # or a release dir / checkpoint
wav = torch.randn(1, 1, 3 * 16000) # [B, 1, samples] at 16 kHz
wav = wav / wav.pow(2).mean().sqrt() # unit RMS per clip, as in training
win = slice_windows(wav, 16000, 120, 80) # [B, n_windows, 1920]
b, n, _ = win.shape
logits_ph, logits_grp = ext.predict(win.reshape(b * n, -1))
stride_frames = 1280 // ext.frame_stride_samples
ph = stitch_windows(logits_ph.reshape(b, n, *logits_ph.shape[1:]), n,
ext.frames_per_window, stride_frames) # [B, frames, 67], 5 ms frames
Forced alignment over a known ph66 sequence is p3cupe.loss.PeakLoss.align with the
config.json decoder settings, as p3cupe/timit_probe.py does.
Training checkpoint
ckpt/en_libri500_pa07f_e0_timit_f1_25ms=75.58.ckpt is the PyTorch Lightning checkpoint (weights only; a pickle, so load it only
if you trust this repo), for ckpt_path: in p3cupe/cupe_hp.yaml or trunk_init_path: in
p4mbfa.
License: AGPL-3.0.
- Downloads last month
- -