|
Download README.md from nyralabs/PINT: direct link, hf CLI and curl.
- Browser
- Download file 2.31 kB
-
https://huggingface.co/nyralabs/PINT/resolve/main/README.md
- Command line
-
hf download hf://nyralabs/PINT/README.md
-
curl -L -o README.md https://huggingface.co/nyralabs/PINT/resolve/main/README.md
2.31 kB
| license: cc-by-nc-4.0 | |
| library_name: pint-infer | |
| pipeline_tag: feature-extraction | |
| base_model: facebook/hubert-base-ls960 | |
| tags: | |
| - speech | |
| - audio | |
| - tokenizer | |
| - hubert | |
| - speech-tokens | |
| language: | |
| - en | |
| # PINT | |
| **Content is What Remains: Invariant Speech Tokenization from Parallel Utterances** — | |
| [arXiv:2607.19033](https://arxiv.org/abs/2607.19033), accepted at **Interspeech 2026**. | |
| A HuBERT-base encoder fine-tuned on parallel utterances (the same words from different speakers and recording conditions) so that they map to the same frame-level representation. 16 kHz audio in, one 768-d vector per 20 ms frame out; token ids come from k-means codebooks. | |
| ## Usage | |
| ```bash | |
| pip install git+https://github.com/nyrahealth/PINT | |
| ``` | |
| ```python | |
| from pint_infer import PINTTokenizer | |
| tok = PINTTokenizer.from_pretrained("nyralabs/PINT") | |
| ids = tok.encode("speech.wav", method="kmeans_200") # one token id per 20 ms | |
| frames = tok.embed("speech.wav") # (T, 768) float32 | |
| ``` | |
| Methods: `kmeans_50`, `kmeans_100`, `kmeans_200`, `kmeans_250`, `kmeans_500`, `kmeans_1000`. | |
| ## Files | |
| | File | Content | | |
| |---|---| | |
| | `model.safetensors`, `config.json`, `preprocessor_config.json` | HuBERT-base encoder and feature extractor | | |
| | `kmeans_<k>.safetensors` | `(k, 768)` k-means cluster centers; assignment is nearest centroid | | |
| | `pint_config.json` | frame rate, codebook list | | |
| k=200 is the codebook size the paper evaluates; the other codebook sizes are provided for convenience. | |
| ## Intended use and limitations | |
| English speech. Input is converted to 16 kHz mono. Research and non-commercial use. | |
| ## Citation | |
| ```bibtex | |
| @misc{wagner2026contentremainsinvariantspeech, | |
| title = {Content is What Remains: Invariant Speech Tokenization from Parallel Utterances}, | |
| author = {Laurin Wagner and Bernhard Thallinger and Miroslav Stankovic and Mario Zusag}, | |
| year = {2026}, | |
| eprint = {2607.19033}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CL}, | |
| url = {https://arxiv.org/abs/2607.19033}, | |
| note = {Accepted at Interspeech 2026} | |
| } | |
| ``` | |
| ## License | |
| The weights are released under CC BY-NC 4.0. They are derived from | |
| [facebook/hubert-base-ls960](https://huggingface.co/facebook/hubert-base-ls960), licensed under | |
| the Apache License 2.0; see `NOTICE`. | |