Tigre HuBERT

A self-supervised speech representation model for Tigre (ISO 639-3: tig), a Semitic language spoken primarily in Eritrea. This is, to our knowledge, one of the first publicly available speech models for Tigre.

Status: this is a pretrained representation model, not a fine-tuned task model. It hasn't yet been benchmarked on a downstream task (e.g. ASR word/character error rate). Validation loss during pretraining improved consistently and has plateaued (see Training below), which is a good sign for representation quality, but the honest way to know how well it performs on a specific task is to fine-tune it for that task and measure the result directly. If you do that, we'd love to hear how it goes.

What this model is

HuBERT (Hidden-Unit BERT) learns speech representations by predicting cluster assignments of masked audio segments, without needing any transcribed text. That makes it well-suited to low-resource languages like Tigre, where labeled speech-to-text data is scarce but raw audio is more attainable.

The output of this model is not text β€” it's a sequence of learned embeddings per audio frame. To get something task-specific (e.g. speech recognition, speaker identification, language identification), you fine-tune a small task head on top of these representations.

How it was trained

  • Data: ~500 hours of Tigre speech audio.
  • Recipe: the standard two-iteration HuBERT pretraining procedure (fairseq's implementation):
    • Iteration 1: pseudo-labels from k-means (k=100) on raw MFCC features β€” a weak but label-free starting signal.
    • Iteration 2: the iteration-1 model's own layer-6 hidden features are re-clustered (k=500) into new, better pseudo-labels, and a fresh model is trained from scratch against those. This is the standard recipe's main quality jump, since the model now learns from speech-aware clusters instead of raw acoustic ones.
  • Architecture: HuBERT-base (~90M parameters, 12 transformer layers).
  • Note on scale: the original HuBERT paper validated this exact 2-iteration recipe on 960 hours of English speech, and only used a 3rd iteration at a much larger scale (60,000+ hours). At ~500 hours, this model is trained at roughly half the data scale the base recipe was designed for β€” a reasonable and appropriately-sized recipe for this amount of data, but worth keeping in mind when setting quality expectations relative to large, high-resource-language HuBERT models.

Iteration-2 training curve

Validation loss over iteration-2 pretraining

Validation loss (the masked-unit prediction objective, not a downstream task metric) over iteration-2 pretraining, from epoch 60 to the point training was stopped once improvement plateaued (under 1% change over the last few checkpoints). This shows the pretraining objective converging smoothly β€” it does not measure accuracy on any specific downstream task.

Basic usage

This model outputs frame-level embeddings, not text. Example: extracting features from a 16kHz mono audio clip.

import torch
from fairseq import checkpoint_utils

# Download checkpoint_best.pt from this repo first
models, cfg, task = checkpoint_utils.load_model_ensemble_and_task(["checkpoint_best.pt"])
model = models[0].eval()

# waveform: torch.FloatTensor of shape [1, num_samples], 16kHz, mono
with torch.inference_mode():
    features, _ = model.extract_features(
        source=waveform,
        padding_mask=None,
        mask=False,       # inference, not the masked-training objective
        output_layer=None # final transformer layer
    )

# features: [1, num_frames, hidden_dim] -- one embedding per ~20ms of audio

fairseq isn't a well-maintained PyPI package β€” install it from source:

git clone https://github.com/facebookresearch/fairseq.git
cd fairseq && pip install --editable .

What this is useful for

  • Fine-tuning for Tigre ASR β€” attach a CTC head and fine-tune on a (even fairly small) labeled Tigre speech-to-text dataset.
  • Speaker or language identification β€” the embeddings can feed a lightweight classifier for tasks that don't need text transcription at all.
  • A starting point, not an endpoint β€” as a foundation model for further Tigre speech research, in a language with very little existing tooling.

Limitations

  • Not evaluated on any downstream task yet β€” treat performance claims with appropriate skepticism until you've tested it on your own task/data.
  • Trained on ~500 hours, notably less than the 960 hours the base recipe was designed around β€” representation quality may reflect that.
  • No fine-tuned ASR head is included in this repo; this model produces embeddings, not transcriptions, out of the box.

Citation

If you use this model, please cite the original HuBERT paper:

@article{hsu2021hubert,
  title={HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units},
  author={Hsu, Wei-Ning and Bolte, Benjamin and Tsai, Yao-Hung Hubert and Lakhotia, Kushal and Salakhutdinov, Ruslan and Mohamed, Abdelrahman},
  journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
  year={2021}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for BeitTigreAI/tigre-hubert