Instructions to use mudler/parakeet-cpp-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use mudler/parakeet-cpp-gguf with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mudler/parakeet-cpp-gguf") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Parakeet GGUF โ models for parakeet.cpp
GGUF-format weights for parakeet.cpp, a C++/ggml port of NVIDIA NeMo Parakeet that matches the upstream PyTorch models on CPU. This single repo collects every supported model ร quantization as a flat set of .gguf files โ download just the one you need.
F16 is the recommended default โ same accuracy as F32, ~1.7ร smaller, and typically the fastest on modern CPUs via ggml's F32รF16 matmul fast path.
Models
tdt_ctc-110m
Source: nvidia/parakeet-tdt_ctc-110m ยท Hybrid TDT+CTC (FastConformer) ยท heads: TDT + CTC
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt_ctc-110m-f16.gguf โ recommended |
F16 | 267.5 MB | 0.0000 |
tdt_ctc-110m-q8_0.gguf |
Q8_0 | 177.8 MB | 0.0000 |
tdt_ctc-110m-q6_k.gguf |
Q6_K | 155.9 MB | not measured |
tdt_ctc-110m-q5_k.gguf |
Q5_K | 143.3 MB | not measured |
tdt_ctc-110m-q4_k.gguf |
Q4_K | 131.4 MB | 0.0000 |
realtime_eou_120m-v1
Source: nvidia/parakeet_realtime_eou_120m-v1 ยท Cache-aware streaming RNNT (FastConformer, EOU/EOB) ยท heads: RNNT (streaming) ยท License: NVIDIA Open Model License
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
realtime_eou_120m-v1-f16.gguf โ recommended |
F16 | 266.5 MB | not measured |
realtime_eou_120m-v1-q8_0.gguf |
Q8_0 | 176.0 MB | not measured |
realtime_eou_120m-v1-q6_k.gguf |
Q6_K | 153.9 MB | not measured |
realtime_eou_120m-v1-q5_k.gguf |
Q5_K | 141.2 MB | not measured |
realtime_eou_120m-v1-q4_k.gguf |
Q4_K | 129.1 MB | not measured |
ctc-0.6b
Source: nvidia/parakeet-ctc-0.6b ยท CTC (FastConformer) ยท heads: CTC
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
ctc-0.6b-f16.gguf โ recommended |
F16 | 1373.4 MB | 0.0000 |
ctc-0.6b-q8_0.gguf |
Q8_0 | 875.4 MB | 0.0000 |
ctc-0.6b-q6_k.gguf |
Q6_K | 746.8 MB | not measured |
ctc-0.6b-q5_k.gguf |
Q5_K | 676.3 MB | not measured |
ctc-0.6b-q4_k.gguf |
Q4_K | 609.9 MB | not measured |
rnnt-0.6b
Source: nvidia/parakeet-rnnt-0.6b ยท RNNT transducer (FastConformer) ยท heads: RNNT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
rnnt-0.6b-f16.gguf โ recommended |
F16 | 1402.8 MB | 0.0000 |
rnnt-0.6b-q8_0.gguf |
Q8_0 | 903.9 MB | 0.0000 |
rnnt-0.6b-q6_k.gguf |
Q6_K | 776.3 MB | not measured |
rnnt-0.6b-q5_k.gguf |
Q5_K | 705.7 MB | not measured |
rnnt-0.6b-q4_k.gguf |
Q4_K | 639.2 MB | not measured |
tdt-0.6b-v2
Source: nvidia/parakeet-tdt-0.6b-v2 ยท TDT transducer (FastConformer) ยท heads: TDT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt-0.6b-v2-f16.gguf โ recommended |
F16 | 1404.2 MB | 0.0000 |
tdt-0.6b-v2-q8_0.gguf |
Q8_0 | 903.8 MB | 0.0000 |
tdt-0.6b-v2-q6_k.gguf |
Q6_K | 775.9 MB | not measured |
tdt-0.6b-v2-q5_k.gguf |
Q5_K | 705.0 MB | not measured |
tdt-0.6b-v2-q4_k.gguf |
Q4_K | 638.4 MB | not measured |
tdt-0.6b-v3
Source: nvidia/parakeet-tdt-0.6b-v3 ยท TDT transducer (FastConformer) ยท heads: TDT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt-0.6b-v3-f16.gguf โ recommended |
F16 | 1441.0 MB | 0.0000 |
tdt-0.6b-v3-q8_0.gguf |
Q8_0 | 940.7 MB | 0.0000 |
tdt-0.6b-v3-q6_k.gguf |
Q6_K | 812.7 MB | not measured |
tdt-0.6b-v3-q5_k.gguf |
Q5_K | 741.9 MB | not measured |
tdt-0.6b-v3-q4_k.gguf |
Q4_K | 675.2 MB | not measured |
ctc-1.1b
Source: nvidia/parakeet-ctc-1.1b ยท CTC (FastConformer) ยท heads: CTC
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
ctc-1.1b-f16.gguf โ recommended |
F16 | 2395.8 MB | 0.0000 |
ctc-1.1b-q8_0.gguf |
Q8_0 | 1526.3 MB | 0.0000 |
ctc-1.1b-q6_k.gguf |
Q6_K | 1301.7 MB | not measured |
ctc-1.1b-q5_k.gguf |
Q5_K | 1178.5 MB | not measured |
ctc-1.1b-q4_k.gguf |
Q4_K | 1062.6 MB | not measured |
rnnt-1.1b
Source: nvidia/parakeet-rnnt-1.1b ยท RNNT transducer (FastConformer) ยท heads: RNNT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
rnnt-1.1b-f16.gguf โ recommended |
F16 | 2425.2 MB | 0.0000 |
rnnt-1.1b-q8_0.gguf |
Q8_0 | 1554.7 MB | 0.0000 |
rnnt-1.1b-q6_k.gguf |
Q6_K | 1331.2 MB | not measured |
rnnt-1.1b-q5_k.gguf |
Q5_K | 1207.9 MB | not measured |
rnnt-1.1b-q4_k.gguf |
Q4_K | 1091.9 MB | not measured |
tdt-1.1b
Source: nvidia/parakeet-tdt-1.1b ยท TDT transducer (FastConformer) ยท heads: TDT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt-1.1b-f16.gguf โ recommended |
F16 | 2425.3 MB | 0.0000 |
tdt-1.1b-q8_0.gguf |
Q8_0 | 1554.8 MB | 0.0000 |
tdt-1.1b-q6_k.gguf |
Q6_K | 1331.2 MB | not measured |
tdt-1.1b-q5_k.gguf |
Q5_K | 1207.9 MB | not measured |
tdt-1.1b-q4_k.gguf |
Q4_K | 1091.9 MB | not measured |
tdt_ctc-1.1b
Source: nvidia/parakeet-tdt_ctc-1.1b ยท Hybrid TDT+CTC (FastConformer) ยท heads: TDT + CTC
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt_ctc-1.1b-f16.gguf โ recommended |
F16 | 2429.5 MB | 0.0000 |
tdt_ctc-1.1b-q8_0.gguf |
Q8_0 | 1559.0 MB | 0.0000 |
tdt_ctc-1.1b-q6_k.gguf |
Q6_K | 1335.4 MB | not measured |
tdt_ctc-1.1b-q5_k.gguf |
Q5_K | 1212.1 MB | not measured |
tdt_ctc-1.1b-q4_k.gguf |
Q4_K | 1096.1 MB | not measured |
WER (word error rate) is computed against the upstream NeMo reference on
tests/fixtures/speech.wav(LibriSpeech2086-149220-0033, ~7.4 s, English). 0.0 = byte-for-byte identical transcript. See parity.md and quantization.md.
nemotron-3.5-asr-streaming-0.6b
Source: nvidia/nemotron-3.5-asr-streaming-0.6b ยท Cache-aware streaming RNNT (FastConformer, 24 encoder layers), multilingual (40 language locales), conditioned on a language prompt ยท heads: RNNT (streaming, prompt-conditioned) ยท License: OpenMDW 1.1
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
nemotron-3.5-asr-streaming-0.6b-f16.gguf โ recommended |
F16 | 1484.3 MB | 0.0000 |
nemotron-3.5-asr-streaming-0.6b-q8_0.gguf |
Q8_0 | 983.7 MB | 0.0000 |
nemotron-3.5-asr-streaming-0.6b-q6_k.gguf |
Q6_K | 855.7 MB | not measured |
nemotron-3.5-asr-streaming-0.6b-q5_k.gguf |
Q5_K | 784.8 MB | not measured |
nemotron-3.5-asr-streaming-0.6b-q4_k.gguf |
Q4_K | 718.1 MB | not measured |
This model runs both offline and as a cache-aware streaming model, and it is the only one here that takes a target language. Pass
--lang <locale>(for exampleen-US,de-DE,es-ES,ja-JP), or leave it at the defaultautoto let the model detect the language. The WER column is agreement with NeMo's transcript on the two short test clips in the parakeet.cpp repository (speech.wav,clip.wav), not accuracy on a speech corpus. The F32 conversion was compared with NeMo for en, de, es, ja-JP and auto, offline and streaming, and matched every time. Q8_0 was compared onspeech.wavin English, offline. The Q6_K, Q5_K and Q4_K files have not been compared yet. See parity.md. The small prompt layers, the LSTM and the feature extractor stay F32 in every quantization.
huggingface-cli download mudler/parakeet-cpp-gguf nemotron-3.5-asr-streaming-0.6b-f16.gguf --local-dir models/
build/examples/cli/parakeet-cli transcribe --model models/nemotron-3.5-asr-streaming-0.6b-f16.gguf --input audio.wav --lang de-DE
nemotron-3-diarization
Source: nvidia/Nemotron-3-Diarization ยท Speaker diarization (Sortformer, up to 8 speakers, streaming speaker cache) ยท License: OpenMDW 1.1
| File | Variant | Size | Segments vs NeMo |
|---|---|---|---|
nemotron-3-diarization-f16.gguf โ recommended |
F16 | 200.7 MB | identical |
nemotron-3-diarization-q8_0.gguf |
Q8_0 | 108.7 MB | identical on 2 of 3 clips, 99.5% of frames on the third |
Checked against NeMo on a 23.6 s and a 68.5 s two-speaker clip (offline and streaming: every segment identical, to the 10 ms frame) and a 12.3 min three-speaker clip (F16 100%, Q8_0 99.9% of speech frames). Q8_0 can flip a frame whose probability sits right at the 0.5 threshold, splitting a segment (99.5% of frames on a 31.5 s clip); use F16 when segment-exact output matters. Answers "who spoke when"; pair it with any ASR model above for speaker-attributed transcripts. See diarization.md.
build/examples/cli/diarize models/nemotron-3-diarization-f16.gguf meeting.wav
ultra and redux (Moondream)
Source: moondream/parakeet-ultra and moondream/parakeet-redux by Moondream, derived from NVIDIA's nvidia/parakeet-tdt-0.6b-v3 ยท TDT transducer (FastConformer), multilingual, with a voice-activity head (transcribe --vad) ยท heads: TDT ยท license: CC-BY-4.0
| File | Variant | Size | Runs on |
|---|---|---|---|
ultra-f16.gguf |
F16 | 1441.9 MB | any backend |
ultra-q8_0.gguf |
Q8_0 | 941.5 MB | any backend |
redux-packed.gguf |
packed ternary encoder | 213.3 MB | CPU only, offline only |
redux-f16.gguf |
F16, dequantized | 1441.9 MB | any backend |
redux-q8_0.gguf |
Q8_0, dequantized | 941.5 MB | any backend |
There is no NeMo reference for these models, so there is no WER column. Each file was checked to decode the
tests/fixtures/speech.wavclip to the expected sentence. See ternary.md for the measurements.
- Ultra has ordinary F16 weights and runs on any backend, like the other v3 files.
- Redux packed keeps the encoder linear layers as ternary weights (-1, 0 or +1 times a per-group scale) and runs a native CPU kernel. It does not run on GPU backends or in streaming mode, and parakeet.cpp refuses to load it there. For GPU use, take
redux-f16.gguforredux-q8_0.gguf. - Redux F16 and Q8_0 are dequantized: the ternary weights were expanded to ordinary weights and then stored as F16 or Q8_0. They are not packed.
- Changes: these files are converted here, not trained. Nothing was trained or fine-tuned by the parakeet.cpp project. The models were trained by NVIDIA (the base) and Moondream (Ultra and Redux).
huggingface-cli download mudler/parakeet-cpp-gguf redux-packed.gguf --local-dir models/
build/examples/cli/parakeet-cli transcribe --model models/redux-packed.gguf --input audio.wav
# Long audio: cut at pauses with the model's voice-activity head (offline only).
build/examples/cli/parakeet-cli transcribe --model models/ultra-q8_0.gguf --input long.wav --vad
silero-vad (Silero Team)
Source: snakers4/silero-vad v6.2.3 by the Silero Team ยท a small voice-activity detector, not a speech recognizer ยท 8 kHz and 16 kHz in one file ยท license: MIT
| File | Variant | Size |
|---|---|---|
silero-vad-f32.gguf |
F32 | 2.2 MB |
silero-vad-f16.gguf |
F16 weights, widened to F32 at load | 1.3 MB |
- The files are converted from the official ONNX model with
scripts/convert_silero_vad_to_gguf.pyin parakeet.cpp. The GGUF records the source version and the ONNX sha256. They were converted here, not trained. - They need a parakeet.cpp build that has the standalone VAD API. That code is not in a release yet, so check the parakeet.cpp repository before relying on it.
- Use it as a stand-alone detector (
parakeet-cli vad --model silero-vad-f16.gguf --input audio.wav) or to cut long audio before transcribing with any model, including those that have no VAD head (parakeet-cli transcribe --model tdt-0.6b-v3-q8_0.gguf --input long.wav --vad --vad-model silero-vad-f16.gguf). - Probabilities match onnxruntime to about 1e-6 (F32) and 1e-3 (F16) on the test clips.
vad heads (Moondream)
Source: the voice-activity head, the mel front end and the subsampler of moondream/parakeet-redux and moondream/parakeet-ultra by Moondream, derived from NVIDIA's nvidia/parakeet-tdt-0.6b-v3 ยท VAD only: these files cannot transcribe ยท license: CC-BY-4.0
| File | Parent | Size |
|---|---|---|
redux-vad.gguf |
redux (the head and subsampler are not ternary) | 9.9 MB |
ultra-vad-q8_0.gguf |
ultra-q8_0 | 6.0 MB |
- Each file holds 20 tensors copied byte for byte from its parent, with no requantization, and records the parent file, its size and its sha256 in the metadata. They were cut out here, not trained.
- Output is identical to the full parent file:
parakeet-cli vad --probabilitiesgave byte-identical JSON on a speech clip, a noisy clip and a 600 s talk. - Against the 213 MB to 941 MB parents, load time falls from about 0.1 to 0.4 s to a few milliseconds, and memory for a 33 s clip from 0.6 to 1.1 GiB to about 0.25 GiB. Speed is the same as the parent's head.
- They need a parakeet.cpp build that can load a VAD-only file. That code is not in a release yet, so check the parakeet.cpp repository before relying on it. Use
parakeet-cli vad --model redux-vad.gguf --input audio.wav. - For a stand-alone detector Silero (above) is smaller and faster per core. The head is useful when you want its recall or already work with the Moondream models.
Bundles
A bundle is one GGUF file that holds several of the models above, so you pass one file instead of several. A bundle can hold an ASR model, a Silero VAD, speaker diarization, sound-event tagging (CED) and speaker identification. A bundle is only packaging: each model is copied byte for byte with the type it was published with, and nothing was re-quantized, trained or fine-tuned. These files are converted here, not trained. The format is described in bundle.md.
| File | Size | Contents | Runs on |
|---|---|---|---|
parakeet-bundle-small.gguf |
337.9 MB | tdt_ctc-110m Q8_0, Nemotron-3-Diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16 | any backend |
parakeet-bundle-standard.gguf |
1100.8 MB | tdt-0.6b-v3 Q8_0, Nemotron-3-Diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16 | any backend |
parakeet-bundle-moondream-redux.gguf |
214.6 MB | redux-packed (with its own VAD head), Silero VAD F16 | CPU only, offline only |
A bundle needs a parakeet.cpp build that includes the bundle code (pull request 85). That code is on the
masterbranch but not in a release yet (the latest release, v0.5.0, does not have it), so build from source for now. Older builds refuse a bundle with a load error. They never read the wrong weights. The bundles were checked on CPU only: the output of each component equals the output of its single-model file (the transcript, the diarization segments, the CED class scores, the speaker embeddings and the Silero probabilities). GPU backends, streaming ASR from a bundle, macOS and Windows were not tested.
huggingface-cli download mudler/parakeet-cpp-gguf parakeet-bundle-small.gguf --local-dir models/
# List the components, licences and credits (reads only the header)
build/examples/cli/parakeet-cli info models/parakeet-bundle-small.gguf
# Transcribe with the ASR component; add --vad to cut long audio with the Silero component
build/examples/cli/parakeet-cli transcribe --model models/parakeet-bundle-small.gguf --input audio.wav
# Who spoke when, with the diarization component
build/examples/cli/diarize models/parakeet-bundle-small.gguf meeting.wav
# Speaker-attributed transcript with sound events: one file passed for every role
build/examples/cli/parakeet-cli scene --model models/parakeet-bundle-small.gguf \
--diar models/parakeet-bundle-small.gguf --sound models/parakeet-bundle-small.gguf \
--input meeting.wav
# Add --speakers models/parakeet-bundle-small.gguf --registry people.bin to name known voices
The moondream-redux bundle has the ASR and Silero components only. Use parakeet-cli transcribe --model models/parakeet-bundle-moondream-redux.gguf --input audio.wav --vad: it cuts at pauses with the Silero component, and --vad-component asr uses the Redux head instead. When a bundle has more than one component of a kind, name one with --component, --asr-component, --diar-component, --sound-component or --speakers-component.
Licences. A bundle has no single licence, so the header says other and each component keeps the licence of the model it was converted from. The credit, the licence link and the changes are in the file header (parakeet-cli info shows them), in NOTICE-parakeet-bundle-<name>.txt next to each bundle, and the full licence texts are in the licenses/ folder of this repo. Keep these notices when you redistribute a bundle.
| Component | Model | Licence | Credit |
|---|---|---|---|
| ASR (small) | nvidia/parakeet-tdt_ctc-110m | CC-BY-4.0 | NVIDIA |
| ASR (standard) | nvidia/parakeet-tdt-0.6b-v3 | CC-BY-4.0 | NVIDIA |
| ASR (moondream-redux) | moondream/parakeet-redux, derived from parakeet-tdt-0.6b-v3 | CC-BY-4.0 | Moondream and NVIDIA |
| Diarization | nvidia/Nemotron-3-Diarization | OpenMDW-1.1 | NVIDIA |
| Sound events | mispeech/ced-small | Apache-2.0 (see the note below) | Heinrich Dinkel et al., Xiaomi (mispeech) |
| Speaker identification | Wespeaker/wespeaker-voxceleb-resnet34-LM | CC-BY-4.0 (see the note below) | the WeSpeaker project |
| VAD | snakers4/silero-vad | MIT | Copyright (c) 2020-present Silero Team |
- CED: the
mispeech/ced-*model cards say Apache-2.0, and the bundle follows them. The upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream. It has not been confirmed with the authors. The CC-BY credit to the authors is kept in the meantime. - WeSpeaker: the file is converted from
voxceleb_resnet34_LM.onnxofWespeaker/wespeaker-voxceleb-resnet34-LM, whose card says CC-BY-4.0. The card of the plainwespeaker-voxceleb-resnet34says Apache-2.0, but the WeSpeaker project states in its documentation that its VoxCeleb-trained models follow CC-BY-4.0, so the bundle uses CC-BY-4.0 and credits the WeSpeaker project. The speaker models are trained on VoxCeleb. Whether a trained model is derived from its training data is a legal question that this project does not settle. - Changes: the models are converted to GGUF here and, for the ASR and diarization models, quantized to Q8_0 (the redux-packed ASR component is the published packed file). Nothing was trained or fine-tuned.
- The end-of-utterance model and the audeering voice-analysis heads are never put in a bundle: their licences do not allow it.
Quantization notes
Quantization is applied only to the large linear weights fed directly into ggml_mul_mat (encoder FFN + attention projections, subsampling output projection, joint enc/pred projections). All other tensors (mel filterbank, LSTM prediction net, conv kernels, batch_norm stats, norms, biases, embeddings) stay F32.
Usage
# 1. Clone + build parakeet.cpp
git clone https://github.com/mudler/parakeet.cpp
cd parakeet.cpp
cmake -B build -DPARAKEET_BUILD_CLI=ON && cmake --build build -j
# 2. Download one quant (F16 recommended)
huggingface-cli download mudler/parakeet-cpp-gguf tdt_ctc-110m-f16.gguf --local-dir models/
# 3. Transcribe
build/examples/cli/parakeet-cli transcribe \
--model models/tdt_ctc-110m-f16.gguf \
--input audio.wav
License
Licences differ by model, so the front matter says license: other. Each file family follows the licence of the model it was converted from:
tdt_ctc-110m-*,ctc-*,rnnt-*,tdt-*andtdt_ctc-1.1b-*: derived from NVIDIA NeMo Parakeet checkpoints released under CC-BY-4.0.realtime_eou_120m-v1-*: derived from nvidia/parakeet_realtime_eou_120m-v1, governed by the NVIDIA Open Model License.nemotron-3.5-asr-streaming-0.6b-*andnemotron-3-diarization-*: derived from NVIDIA Nemotron models, governed by the OpenMDW License Agreement, version 1.1.ultra-*andredux-*: see below.parakeet-bundle-*: one licence per component, see Bundles.
The notes below add detail. ultra-*.gguf and redux-*.gguf are converted from moondream/parakeet-ultra and moondream/parakeet-redux by Moondream, which are derived from NVIDIA's parakeet-tdt-0.6b-v3. Both are also CC-BY-4.0: credit Moondream and NVIDIA when you use these files. They were converted here, not trained, and the Redux F16 and Q8_0 files are dequantized from the ternary weights. redux-vad.gguf and ultra-vad-q8_0.gguf hold only the VAD head, front end and subsampler of the Moondream models, cut out of the files above under the same CC-BY-4.0 terms: credit Moondream and NVIDIA; they were cut out here, not trained. silero-vad-*.gguf is converted from Silero VAD v6.2.3 and is released under the MIT license, Copyright (c) 2020-present Silero Team; it was converted here, not trained. nemotron-3-diarization-*.gguf is derived from nvidia/Nemotron-3-Diarization and nemotron-3.5-asr-streaming-0.6b-*.gguf from nvidia/nemotron-3.5-asr-streaming-0.6b. Both are governed by the OpenMDW License Agreement, version 1.1. The parakeet.cpp runtime is MIT-licensed.
- Downloads last month
- 23,627
Model tree for mudler/parakeet-cpp-gguf
Base model
Wespeaker/wespeaker-voxceleb-resnet34-LM
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mudler/parakeet-cpp-gguf") transcriptions = asr_model.transcribe(["file.wav"])