--- license: cc-by-nc-4.0 library_name: pint-infer pipeline_tag: feature-extraction base_model: facebook/hubert-base-ls960 tags: - speech - audio - tokenizer - hubert - speech-tokens language: - en --- # PINT **Content is What Remains: Invariant Speech Tokenization from Parallel Utterances** — [arXiv:2607.19033](https://arxiv.org/abs/2607.19033), accepted at **Interspeech 2026**. A HuBERT-base encoder fine-tuned on parallel utterances (the same words from different speakers and recording conditions) so that they map to the same frame-level representation. 16 kHz audio in, one 768-d vector per 20 ms frame out; token ids come from k-means codebooks. ## Usage ```bash pip install git+https://github.com/nyrahealth/PINT ``` ```python from pint_infer import PINTTokenizer tok = PINTTokenizer.from_pretrained("nyralabs/PINT") ids = tok.encode("speech.wav", method="kmeans_200") # one token id per 20 ms frames = tok.embed("speech.wav") # (T, 768) float32 ``` Methods: `kmeans_50`, `kmeans_100`, `kmeans_200`, `kmeans_250`, `kmeans_500`, `kmeans_1000`. ## Files | File | Content | |---|---| | `model.safetensors`, `config.json`, `preprocessor_config.json` | HuBERT-base encoder and feature extractor | | `kmeans_.safetensors` | `(k, 768)` k-means cluster centers; assignment is nearest centroid | | `pint_config.json` | frame rate, codebook list | k=200 is the codebook size the paper evaluates; the other codebook sizes are provided for convenience. ## Intended use and limitations English speech. Input is converted to 16 kHz mono. Research and non-commercial use. ## Citation ```bibtex @misc{wagner2026contentremainsinvariantspeech, title = {Content is What Remains: Invariant Speech Tokenization from Parallel Utterances}, author = {Laurin Wagner and Bernhard Thallinger and Miroslav Stankovic and Mario Zusag}, year = {2026}, eprint = {2607.19033}, archivePrefix = {arXiv}, primaryClass = {cs.CL}, url = {https://arxiv.org/abs/2607.19033}, note = {Accepted at Interspeech 2026} } ``` ## License The weights are released under CC BY-NC 4.0. They are derived from [facebook/hubert-base-ls960](https://huggingface.co/facebook/hubert-base-ls960), licensed under the Apache License 2.0; see `NOTICE`.