PINT
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances — arXiv:2607.19033, accepted at Interspeech 2026.
A HuBERT-base encoder fine-tuned on parallel utterances (the same words from different speakers and recording conditions) so that they map to the same frame-level representation. 16 kHz audio in, one 768-d vector per 20 ms frame out; token ids come from k-means codebooks.
Usage
pip install git+https://github.com/nyrahealth/PINT
from pint_infer import PINTTokenizer
tok = PINTTokenizer.from_pretrained("nyralabs/PINT")
ids = tok.encode("speech.wav", method="kmeans_200") # one token id per 20 ms
frames = tok.embed("speech.wav") # (T, 768) float32
Methods: kmeans_50, kmeans_100, kmeans_200, kmeans_250, kmeans_500, kmeans_1000.
Files
| File | Content |
|---|---|
model.safetensors, config.json, preprocessor_config.json |
HuBERT-base encoder and feature extractor |
kmeans_<k>.safetensors |
(k, 768) k-means cluster centers; assignment is nearest centroid |
pint_config.json |
frame rate, codebook list |
k=200 is the codebook size the paper evaluates; the other codebook sizes are provided for convenience.
Intended use and limitations
English speech. Input is converted to 16 kHz mono. Research and non-commercial use.
Citation
@misc{wagner2026contentremainsinvariantspeech,
title = {Content is What Remains: Invariant Speech Tokenization from Parallel Utterances},
author = {Laurin Wagner and Bernhard Thallinger and Miroslav Stankovic and Mario Zusag},
year = {2026},
eprint = {2607.19033},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2607.19033},
note = {Accepted at Interspeech 2026}
}
License
The weights are released under CC BY-NC 4.0. They are derived from
facebook/hubert-base-ls960, licensed under
the Apache License 2.0; see NOTICE.
- Downloads last month
- 17
Model tree for nyralabs/PINT
Base model
facebook/hubert-base-ls960