PINT

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances — arXiv:2607.19033, accepted at Interspeech 2026.

A HuBERT-base encoder fine-tuned on parallel utterances (the same words from different speakers and recording conditions) so that they map to the same frame-level representation. 16 kHz audio in, one 768-d vector per 20 ms frame out; token ids come from k-means codebooks.

Usage

pip install git+https://github.com/nyrahealth/PINT
from pint_infer import PINTTokenizer

tok = PINTTokenizer.from_pretrained("nyralabs/PINT")
ids = tok.encode("speech.wav", method="kmeans_200")  # one token id per 20 ms
frames = tok.embed("speech.wav")  # (T, 768) float32

Methods: kmeans_50, kmeans_100, kmeans_200, kmeans_250, kmeans_500, kmeans_1000.

Files

File Content
model.safetensors, config.json, preprocessor_config.json HuBERT-base encoder and feature extractor
kmeans_<k>.safetensors (k, 768) k-means cluster centers; assignment is nearest centroid
pint_config.json frame rate, codebook list

k=200 is the codebook size the paper evaluates; the other codebook sizes are provided for convenience.

Intended use and limitations

English speech. Input is converted to 16 kHz mono. Research and non-commercial use.

Citation

@misc{wagner2026contentremainsinvariantspeech,
  title         = {Content is What Remains: Invariant Speech Tokenization from Parallel Utterances},
  author        = {Laurin Wagner and Bernhard Thallinger and Miroslav Stankovic and Mario Zusag},
  year          = {2026},
  eprint        = {2607.19033},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2607.19033},
  note          = {Accepted at Interspeech 2026}
}

License

The weights are released under CC BY-NC 4.0. They are derived from facebook/hubert-base-ls960, licensed under the Apache License 2.0; see NOTICE.

Downloads last month
17
Safetensors
Model size
94.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nyralabs/PINT

Finetuned
(155)
this model

Paper for nyralabs/PINT