PINT / README.md
btnh's picture
Model card: install from github.com/nyrahealth/PINT
f0fe58f verified
|
Raw History Blame Contribute Delete
2.31 kB
---
license: cc-by-nc-4.0
library_name: pint-infer
pipeline_tag: feature-extraction
base_model: facebook/hubert-base-ls960
tags:
- speech
- audio
- tokenizer
- hubert
- speech-tokens
language:
- en
---
# PINT
**Content is What Remains: Invariant Speech Tokenization from Parallel Utterances** —
[arXiv:2607.19033](https://arxiv.org/abs/2607.19033), accepted at **Interspeech 2026**.
A HuBERT-base encoder fine-tuned on parallel utterances (the same words from different speakers and recording conditions) so that they map to the same frame-level representation. 16 kHz audio in, one 768-d vector per 20 ms frame out; token ids come from k-means codebooks.
## Usage
```bash
pip install git+https://github.com/nyrahealth/PINT
```
```python
from pint_infer import PINTTokenizer
tok = PINTTokenizer.from_pretrained("nyralabs/PINT")
ids = tok.encode("speech.wav", method="kmeans_200") # one token id per 20 ms
frames = tok.embed("speech.wav") # (T, 768) float32
```
Methods: `kmeans_50`, `kmeans_100`, `kmeans_200`, `kmeans_250`, `kmeans_500`, `kmeans_1000`.
## Files
| File | Content |
|---|---|
| `model.safetensors`, `config.json`, `preprocessor_config.json` | HuBERT-base encoder and feature extractor |
| `kmeans_<k>.safetensors` | `(k, 768)` k-means cluster centers; assignment is nearest centroid |
| `pint_config.json` | frame rate, codebook list |
k=200 is the codebook size the paper evaluates; the other codebook sizes are provided for convenience.
## Intended use and limitations
English speech. Input is converted to 16 kHz mono. Research and non-commercial use.
## Citation
```bibtex
@misc{wagner2026contentremainsinvariantspeech,
title = {Content is What Remains: Invariant Speech Tokenization from Parallel Utterances},
author = {Laurin Wagner and Bernhard Thallinger and Miroslav Stankovic and Mario Zusag},
year = {2026},
eprint = {2607.19033},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2607.19033},
note = {Accepted at Interspeech 2026}
}
```
## License
The weights are released under CC BY-NC 4.0. They are derived from
[facebook/hubert-base-ls960](https://huggingface.co/facebook/hubert-base-ls960), licensed under
the Apache License 2.0; see `NOTICE`.