CUPE v3 (p3cupe) 路 English

A context-free CNN phoneme encoder (the third CUPE), trained without boundary labels to place phoneme boundaries: one ph66 phoneme distribution every 5 ms, each frame seeing at most 120 ms of audio (38.9 ms receptive field). Boundaries come from a segmental Viterbi decoder over the known phone sequence, read at the crossing of neighbouring posteriors.

Superseded by p4mbfa (Tabahi/mbfa-<language group>, standard_g2p labels, many languages). Kept for BFA's ph66 pipeline.

head labels size
phonemes ph66 (config.json labels.phonemes, index 0 = SIL; the last slot, 'noise', is never trained) 67
groups ph66 phoneme groups (labels.phoneme_groups) 17

Training

  • Data: LibriSpeech, libri_500h (~497 h), one epoch.
  • Checkpoint: experiment pa07f, epoch 0, selected on TIMIT lev f1@25ms on the tune half, re-scored with the final decoder.

Warm-started from pa07e; the lineage is pa01d -> pa07a (MLS, 7 languages, 100 h) -> pa07b (libri_360h) -> pa07c, pa07d, pa07e (libri_500h) -> pa07f. Trained with min_phone_frames 3, soft 10, onset bonus 30, ramp 3. For per-phone onset accuracy decode with boundary_bonus_weight 0 (phone_hit 73.07 vs 68.73).

metric value
timit_lev_f1_25ms_tune 76.1
timit_lev_f1_25ms_holdout 75.58
timit_lev_f1_25ms_bonus0_holdout 73.77
timit_phone_hit_25ms_bonus0_holdout 73.07
timit_phone_hit_25ms_holdout 68.73

TIMIT: 400 clips (English, never trained on), split in a tune half (picked the checkpoint and the decoder) and a holdout half. lev_f1_25ms is phone-boundary F1 within 25 ms; phone_hit_25ms is the share of phones whose onset is within 25 ms. Scored with the decoder in config.json ("decoder": min 6 + soft 10 frames, log-mel onset bonus 30.0); bonus0 = the same decoder without the bonus. The settings it was trained with are under training.decoder_at_training.

Usage

The code is in https://github.com/tabahi/bfa_models (p3cupe/); clone it and run from its root.

import torch
from p3cupe.model import CupeExtractor
from p3cupe.windowing import slice_windows, stitch_windows

ext = CupeExtractor("Tabahi/CUPE-3-en", device="cpu")         # or a release dir / checkpoint
wav = torch.randn(1, 1, 3 * 16000)                     # [B, 1, samples] at 16 kHz
wav = wav / wav.pow(2).mean().sqrt()                   # unit RMS per clip, as in training
win = slice_windows(wav, 16000, 120, 80)  # [B, n_windows, 1920]
b, n, _ = win.shape
logits_ph, logits_grp = ext.predict(win.reshape(b * n, -1))
stride_frames = 1280 // ext.frame_stride_samples
ph = stitch_windows(logits_ph.reshape(b, n, *logits_ph.shape[1:]), n,
                    ext.frames_per_window, stride_frames)  # [B, frames, 67], 5 ms frames

Forced alignment over a known ph66 sequence is p3cupe.loss.PeakLoss.align with the config.json decoder settings, as p3cupe/timit_probe.py does.

Training checkpoint

ckpt/en_libri500_pa07f_e0_timit_f1_25ms=75.58.ckpt is the PyTorch Lightning checkpoint (weights only; a pickle, so load it only if you trust this repo), for ckpt_path: in p3cupe/cupe_hp.yaml or trunk_init_path: in p4mbfa.

License: AGPL-3.0.

Downloads last month
-
Safetensors
Model size
13.5M params
Tensor type
F32
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support