How to use from the
Use from the
NeMo library
import nemo.collections.asr as nemo_asr
asr_model = nemo_asr.models.ASRModel.from_pretrained("mudler/parakeet-cpp-gguf")

transcriptions = asr_model.transcribe(["file.wav"])

Parakeet GGUF โ€” models for parakeet.cpp

GGUF-format weights for parakeet.cpp, a C++/ggml port of NVIDIA NeMo Parakeet that matches the upstream PyTorch models on CPU. This single repo collects every supported model ร— quantization as a flat set of .gguf files โ€” download just the one you need.

F16 is the recommended default โ€” same accuracy as F32, ~1.7ร— smaller, and typically the fastest on modern CPUs via ggml's F32ร—F16 matmul fast path.

Models

tdt_ctc-110m

Source: nvidia/parakeet-tdt_ctc-110m ยท Hybrid TDT+CTC (FastConformer) ยท heads: TDT + CTC

File Variant Size WER vs NeMo
tdt_ctc-110m-f16.gguf โ† recommended F16 267.5 MB 0.0000
tdt_ctc-110m-q8_0.gguf Q8_0 177.8 MB 0.0000
tdt_ctc-110m-q6_k.gguf Q6_K 155.9 MB not measured
tdt_ctc-110m-q5_k.gguf Q5_K 143.3 MB not measured
tdt_ctc-110m-q4_k.gguf Q4_K 131.4 MB 0.0000

realtime_eou_120m-v1

Source: nvidia/parakeet_realtime_eou_120m-v1 ยท Cache-aware streaming RNNT (FastConformer, EOU/EOB) ยท heads: RNNT (streaming) ยท License: NVIDIA Open Model License

File Variant Size WER vs NeMo
realtime_eou_120m-v1-f16.gguf โ† recommended F16 266.5 MB not measured
realtime_eou_120m-v1-q8_0.gguf Q8_0 176.0 MB not measured
realtime_eou_120m-v1-q6_k.gguf Q6_K 153.9 MB not measured
realtime_eou_120m-v1-q5_k.gguf Q5_K 141.2 MB not measured
realtime_eou_120m-v1-q4_k.gguf Q4_K 129.1 MB not measured

ctc-0.6b

Source: nvidia/parakeet-ctc-0.6b ยท CTC (FastConformer) ยท heads: CTC

File Variant Size WER vs NeMo
ctc-0.6b-f16.gguf โ† recommended F16 1373.4 MB 0.0000
ctc-0.6b-q8_0.gguf Q8_0 875.4 MB 0.0000
ctc-0.6b-q6_k.gguf Q6_K 746.8 MB not measured
ctc-0.6b-q5_k.gguf Q5_K 676.3 MB not measured
ctc-0.6b-q4_k.gguf Q4_K 609.9 MB not measured

rnnt-0.6b

Source: nvidia/parakeet-rnnt-0.6b ยท RNNT transducer (FastConformer) ยท heads: RNNT

File Variant Size WER vs NeMo
rnnt-0.6b-f16.gguf โ† recommended F16 1402.8 MB 0.0000
rnnt-0.6b-q8_0.gguf Q8_0 903.9 MB 0.0000
rnnt-0.6b-q6_k.gguf Q6_K 776.3 MB not measured
rnnt-0.6b-q5_k.gguf Q5_K 705.7 MB not measured
rnnt-0.6b-q4_k.gguf Q4_K 639.2 MB not measured

tdt-0.6b-v2

Source: nvidia/parakeet-tdt-0.6b-v2 ยท TDT transducer (FastConformer) ยท heads: TDT

File Variant Size WER vs NeMo
tdt-0.6b-v2-f16.gguf โ† recommended F16 1404.2 MB 0.0000
tdt-0.6b-v2-q8_0.gguf Q8_0 903.8 MB 0.0000
tdt-0.6b-v2-q6_k.gguf Q6_K 775.9 MB not measured
tdt-0.6b-v2-q5_k.gguf Q5_K 705.0 MB not measured
tdt-0.6b-v2-q4_k.gguf Q4_K 638.4 MB not measured

tdt-0.6b-v3

Source: nvidia/parakeet-tdt-0.6b-v3 ยท TDT transducer (FastConformer) ยท heads: TDT

File Variant Size WER vs NeMo
tdt-0.6b-v3-f16.gguf โ† recommended F16 1441.0 MB 0.0000
tdt-0.6b-v3-q8_0.gguf Q8_0 940.7 MB 0.0000
tdt-0.6b-v3-q6_k.gguf Q6_K 812.7 MB not measured
tdt-0.6b-v3-q5_k.gguf Q5_K 741.9 MB not measured
tdt-0.6b-v3-q4_k.gguf Q4_K 675.2 MB not measured

ctc-1.1b

Source: nvidia/parakeet-ctc-1.1b ยท CTC (FastConformer) ยท heads: CTC

File Variant Size WER vs NeMo
ctc-1.1b-f16.gguf โ† recommended F16 2395.8 MB 0.0000
ctc-1.1b-q8_0.gguf Q8_0 1526.3 MB 0.0000
ctc-1.1b-q6_k.gguf Q6_K 1301.7 MB not measured
ctc-1.1b-q5_k.gguf Q5_K 1178.5 MB not measured
ctc-1.1b-q4_k.gguf Q4_K 1062.6 MB not measured

rnnt-1.1b

Source: nvidia/parakeet-rnnt-1.1b ยท RNNT transducer (FastConformer) ยท heads: RNNT

File Variant Size WER vs NeMo
rnnt-1.1b-f16.gguf โ† recommended F16 2425.2 MB 0.0000
rnnt-1.1b-q8_0.gguf Q8_0 1554.7 MB 0.0000
rnnt-1.1b-q6_k.gguf Q6_K 1331.2 MB not measured
rnnt-1.1b-q5_k.gguf Q5_K 1207.9 MB not measured
rnnt-1.1b-q4_k.gguf Q4_K 1091.9 MB not measured

tdt-1.1b

Source: nvidia/parakeet-tdt-1.1b ยท TDT transducer (FastConformer) ยท heads: TDT

File Variant Size WER vs NeMo
tdt-1.1b-f16.gguf โ† recommended F16 2425.3 MB 0.0000
tdt-1.1b-q8_0.gguf Q8_0 1554.8 MB 0.0000
tdt-1.1b-q6_k.gguf Q6_K 1331.2 MB not measured
tdt-1.1b-q5_k.gguf Q5_K 1207.9 MB not measured
tdt-1.1b-q4_k.gguf Q4_K 1091.9 MB not measured

tdt_ctc-1.1b

Source: nvidia/parakeet-tdt_ctc-1.1b ยท Hybrid TDT+CTC (FastConformer) ยท heads: TDT + CTC

File Variant Size WER vs NeMo
tdt_ctc-1.1b-f16.gguf โ† recommended F16 2429.5 MB 0.0000
tdt_ctc-1.1b-q8_0.gguf Q8_0 1559.0 MB 0.0000
tdt_ctc-1.1b-q6_k.gguf Q6_K 1335.4 MB not measured
tdt_ctc-1.1b-q5_k.gguf Q5_K 1212.1 MB not measured
tdt_ctc-1.1b-q4_k.gguf Q4_K 1096.1 MB not measured

WER (word error rate) is computed against the upstream NeMo reference on tests/fixtures/speech.wav (LibriSpeech 2086-149220-0033, ~7.4 s, English). 0.0 = byte-for-byte identical transcript. See parity.md and quantization.md.

nemotron-3.5-asr-streaming-0.6b

Source: nvidia/nemotron-3.5-asr-streaming-0.6b ยท Cache-aware streaming RNNT (FastConformer, 24 encoder layers), multilingual (40 language locales), conditioned on a language prompt ยท heads: RNNT (streaming, prompt-conditioned) ยท License: OpenMDW 1.1

File Variant Size WER vs NeMo
nemotron-3.5-asr-streaming-0.6b-f16.gguf โ† recommended F16 1484.3 MB 0.0000
nemotron-3.5-asr-streaming-0.6b-q8_0.gguf Q8_0 983.7 MB 0.0000
nemotron-3.5-asr-streaming-0.6b-q6_k.gguf Q6_K 855.7 MB not measured
nemotron-3.5-asr-streaming-0.6b-q5_k.gguf Q5_K 784.8 MB not measured
nemotron-3.5-asr-streaming-0.6b-q4_k.gguf Q4_K 718.1 MB not measured

This model runs both offline and as a cache-aware streaming model, and it is the only one here that takes a target language. Pass --lang <locale> (for example en-US, de-DE, es-ES, ja-JP), or leave it at the default auto to let the model detect the language. The WER column is agreement with NeMo's transcript on the two short test clips in the parakeet.cpp repository (speech.wav, clip.wav), not accuracy on a speech corpus. The F32 conversion was compared with NeMo for en, de, es, ja-JP and auto, offline and streaming, and matched every time. Q8_0 was compared on speech.wav in English, offline. The Q6_K, Q5_K and Q4_K files have not been compared yet. See parity.md. The small prompt layers, the LSTM and the feature extractor stay F32 in every quantization.

huggingface-cli download mudler/parakeet-cpp-gguf nemotron-3.5-asr-streaming-0.6b-f16.gguf --local-dir models/
build/examples/cli/parakeet-cli transcribe --model models/nemotron-3.5-asr-streaming-0.6b-f16.gguf --input audio.wav --lang de-DE

nemotron-3-diarization

Source: nvidia/Nemotron-3-Diarization ยท Speaker diarization (Sortformer, up to 8 speakers, streaming speaker cache) ยท License: OpenMDW 1.1

File Variant Size Segments vs NeMo
nemotron-3-diarization-f16.gguf โ† recommended F16 200.7 MB identical
nemotron-3-diarization-q8_0.gguf Q8_0 108.7 MB identical on 2 of 3 clips, 99.5% of frames on the third

Checked against NeMo on a 23.6 s and a 68.5 s two-speaker clip (offline and streaming: every segment identical, to the 10 ms frame) and a 12.3 min three-speaker clip (F16 100%, Q8_0 99.9% of speech frames). Q8_0 can flip a frame whose probability sits right at the 0.5 threshold, splitting a segment (99.5% of frames on a 31.5 s clip); use F16 when segment-exact output matters. Answers "who spoke when"; pair it with any ASR model above for speaker-attributed transcripts. See diarization.md.

build/examples/cli/diarize models/nemotron-3-diarization-f16.gguf meeting.wav

ultra and redux (Moondream)

Source: moondream/parakeet-ultra and moondream/parakeet-redux by Moondream, derived from NVIDIA's nvidia/parakeet-tdt-0.6b-v3 ยท TDT transducer (FastConformer), multilingual, with a voice-activity head (transcribe --vad) ยท heads: TDT ยท license: CC-BY-4.0

File Variant Size Runs on
ultra-f16.gguf F16 1441.9 MB any backend
ultra-q8_0.gguf Q8_0 941.5 MB any backend
redux-packed.gguf packed ternary encoder 213.3 MB CPU only, offline only
redux-f16.gguf F16, dequantized 1441.9 MB any backend
redux-q8_0.gguf Q8_0, dequantized 941.5 MB any backend

There is no NeMo reference for these models, so there is no WER column. Each file was checked to decode the tests/fixtures/speech.wav clip to the expected sentence. See ternary.md for the measurements.

  • Ultra has ordinary F16 weights and runs on any backend, like the other v3 files.
  • Redux packed keeps the encoder linear layers as ternary weights (-1, 0 or +1 times a per-group scale) and runs a native CPU kernel. It does not run on GPU backends or in streaming mode, and parakeet.cpp refuses to load it there. For GPU use, take redux-f16.gguf or redux-q8_0.gguf.
  • Redux F16 and Q8_0 are dequantized: the ternary weights were expanded to ordinary weights and then stored as F16 or Q8_0. They are not packed.
  • Changes: these files are converted here, not trained. Nothing was trained or fine-tuned by the parakeet.cpp project. The models were trained by NVIDIA (the base) and Moondream (Ultra and Redux).
huggingface-cli download mudler/parakeet-cpp-gguf redux-packed.gguf --local-dir models/
build/examples/cli/parakeet-cli transcribe --model models/redux-packed.gguf --input audio.wav
# Long audio: cut at pauses with the model's voice-activity head (offline only).
build/examples/cli/parakeet-cli transcribe --model models/ultra-q8_0.gguf --input long.wav --vad

silero-vad (Silero Team)

Source: snakers4/silero-vad v6.2.3 by the Silero Team ยท a small voice-activity detector, not a speech recognizer ยท 8 kHz and 16 kHz in one file ยท license: MIT

File Variant Size
silero-vad-f32.gguf F32 2.2 MB
silero-vad-f16.gguf F16 weights, widened to F32 at load 1.3 MB
  • The files are converted from the official ONNX model with scripts/convert_silero_vad_to_gguf.py in parakeet.cpp. The GGUF records the source version and the ONNX sha256. They were converted here, not trained.
  • They need a parakeet.cpp build that has the standalone VAD API. That code is not in a release yet, so check the parakeet.cpp repository before relying on it.
  • Use it as a stand-alone detector (parakeet-cli vad --model silero-vad-f16.gguf --input audio.wav) or to cut long audio before transcribing with any model, including those that have no VAD head (parakeet-cli transcribe --model tdt-0.6b-v3-q8_0.gguf --input long.wav --vad --vad-model silero-vad-f16.gguf).
  • Probabilities match onnxruntime to about 1e-6 (F32) and 1e-3 (F16) on the test clips.

vad heads (Moondream)

Source: the voice-activity head, the mel front end and the subsampler of moondream/parakeet-redux and moondream/parakeet-ultra by Moondream, derived from NVIDIA's nvidia/parakeet-tdt-0.6b-v3 ยท VAD only: these files cannot transcribe ยท license: CC-BY-4.0

File Parent Size
redux-vad.gguf redux (the head and subsampler are not ternary) 9.9 MB
ultra-vad-q8_0.gguf ultra-q8_0 6.0 MB
  • Each file holds 20 tensors copied byte for byte from its parent, with no requantization, and records the parent file, its size and its sha256 in the metadata. They were cut out here, not trained.
  • Output is identical to the full parent file: parakeet-cli vad --probabilities gave byte-identical JSON on a speech clip, a noisy clip and a 600 s talk.
  • Against the 213 MB to 941 MB parents, load time falls from about 0.1 to 0.4 s to a few milliseconds, and memory for a 33 s clip from 0.6 to 1.1 GiB to about 0.25 GiB. Speed is the same as the parent's head.
  • They need a parakeet.cpp build that can load a VAD-only file. That code is not in a release yet, so check the parakeet.cpp repository before relying on it. Use parakeet-cli vad --model redux-vad.gguf --input audio.wav.
  • For a stand-alone detector Silero (above) is smaller and faster per core. The head is useful when you want its recall or already work with the Moondream models.

Bundles

A bundle is one GGUF file that holds several of the models above, so you pass one file instead of several. A bundle can hold an ASR model, a Silero VAD, speaker diarization, sound-event tagging (CED) and speaker identification. A bundle is only packaging: each model is copied byte for byte with the type it was published with, and nothing was re-quantized, trained or fine-tuned. These files are converted here, not trained. The format is described in bundle.md.

File Size Contents Runs on
parakeet-bundle-small.gguf 337.9 MB tdt_ctc-110m Q8_0, Nemotron-3-Diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16 any backend
parakeet-bundle-standard.gguf 1100.8 MB tdt-0.6b-v3 Q8_0, Nemotron-3-Diarization Q8_0, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16 any backend
parakeet-bundle-moondream-redux.gguf 214.6 MB redux-packed (with its own VAD head), Silero VAD F16 CPU only, offline only

A bundle needs a parakeet.cpp build that includes the bundle code (pull request 85). That code is on the master branch but not in a release yet (the latest release, v0.5.0, does not have it), so build from source for now. Older builds refuse a bundle with a load error. They never read the wrong weights. The bundles were checked on CPU only: the output of each component equals the output of its single-model file (the transcript, the diarization segments, the CED class scores, the speaker embeddings and the Silero probabilities). GPU backends, streaming ASR from a bundle, macOS and Windows were not tested.

huggingface-cli download mudler/parakeet-cpp-gguf parakeet-bundle-small.gguf --local-dir models/

# List the components, licences and credits (reads only the header)
build/examples/cli/parakeet-cli info models/parakeet-bundle-small.gguf

# Transcribe with the ASR component; add --vad to cut long audio with the Silero component
build/examples/cli/parakeet-cli transcribe --model models/parakeet-bundle-small.gguf --input audio.wav

# Who spoke when, with the diarization component
build/examples/cli/diarize models/parakeet-bundle-small.gguf meeting.wav

# Speaker-attributed transcript with sound events: one file passed for every role
build/examples/cli/parakeet-cli scene --model models/parakeet-bundle-small.gguf \
    --diar models/parakeet-bundle-small.gguf --sound models/parakeet-bundle-small.gguf \
    --input meeting.wav
# Add --speakers models/parakeet-bundle-small.gguf --registry people.bin to name known voices

The moondream-redux bundle has the ASR and Silero components only. Use parakeet-cli transcribe --model models/parakeet-bundle-moondream-redux.gguf --input audio.wav --vad: it cuts at pauses with the Silero component, and --vad-component asr uses the Redux head instead. When a bundle has more than one component of a kind, name one with --component, --asr-component, --diar-component, --sound-component or --speakers-component.

Licences. A bundle has no single licence, so the header says other and each component keeps the licence of the model it was converted from. The credit, the licence link and the changes are in the file header (parakeet-cli info shows them), in NOTICE-parakeet-bundle-<name>.txt next to each bundle, and the full licence texts are in the licenses/ folder of this repo. Keep these notices when you redistribute a bundle.

Component Model Licence Credit
ASR (small) nvidia/parakeet-tdt_ctc-110m CC-BY-4.0 NVIDIA
ASR (standard) nvidia/parakeet-tdt-0.6b-v3 CC-BY-4.0 NVIDIA
ASR (moondream-redux) moondream/parakeet-redux, derived from parakeet-tdt-0.6b-v3 CC-BY-4.0 Moondream and NVIDIA
Diarization nvidia/Nemotron-3-Diarization OpenMDW-1.1 NVIDIA
Sound events mispeech/ced-small Apache-2.0 (see the note below) Heinrich Dinkel et al., Xiaomi (mispeech)
Speaker identification Wespeaker/wespeaker-voxceleb-resnet34-LM CC-BY-4.0 (see the note below) the WeSpeaker project
VAD snakers4/silero-vad MIT Copyright (c) 2020-present Silero Team
  • CED: the mispeech/ced-* model cards say Apache-2.0, and the bundle follows them. The upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream. It has not been confirmed with the authors. The CC-BY credit to the authors is kept in the meantime.
  • WeSpeaker: the file is converted from voxceleb_resnet34_LM.onnx of Wespeaker/wespeaker-voxceleb-resnet34-LM, whose card says CC-BY-4.0. The card of the plain wespeaker-voxceleb-resnet34 says Apache-2.0, but the WeSpeaker project states in its documentation that its VoxCeleb-trained models follow CC-BY-4.0, so the bundle uses CC-BY-4.0 and credits the WeSpeaker project. The speaker models are trained on VoxCeleb. Whether a trained model is derived from its training data is a legal question that this project does not settle.
  • Changes: the models are converted to GGUF here and, for the ASR and diarization models, quantized to Q8_0 (the redux-packed ASR component is the published packed file). Nothing was trained or fine-tuned.
  • The end-of-utterance model and the audeering voice-analysis heads are never put in a bundle: their licences do not allow it.

Quantization notes

Quantization is applied only to the large linear weights fed directly into ggml_mul_mat (encoder FFN + attention projections, subsampling output projection, joint enc/pred projections). All other tensors (mel filterbank, LSTM prediction net, conv kernels, batch_norm stats, norms, biases, embeddings) stay F32.

Usage

# 1. Clone + build parakeet.cpp
git clone https://github.com/mudler/parakeet.cpp
cd parakeet.cpp
cmake -B build -DPARAKEET_BUILD_CLI=ON && cmake --build build -j

# 2. Download one quant (F16 recommended)
huggingface-cli download mudler/parakeet-cpp-gguf tdt_ctc-110m-f16.gguf --local-dir models/

# 3. Transcribe
build/examples/cli/parakeet-cli transcribe \
    --model models/tdt_ctc-110m-f16.gguf \
    --input audio.wav

License

Licences differ by model, so the front matter says license: other. Each file family follows the licence of the model it was converted from:

The notes below add detail. ultra-*.gguf and redux-*.gguf are converted from moondream/parakeet-ultra and moondream/parakeet-redux by Moondream, which are derived from NVIDIA's parakeet-tdt-0.6b-v3. Both are also CC-BY-4.0: credit Moondream and NVIDIA when you use these files. They were converted here, not trained, and the Redux F16 and Q8_0 files are dequantized from the ternary weights. redux-vad.gguf and ultra-vad-q8_0.gguf hold only the VAD head, front end and subsampler of the Moondream models, cut out of the files above under the same CC-BY-4.0 terms: credit Moondream and NVIDIA; they were cut out here, not trained. silero-vad-*.gguf is converted from Silero VAD v6.2.3 and is released under the MIT license, Copyright (c) 2020-present Silero Team; it was converted here, not trained. nemotron-3-diarization-*.gguf is derived from nvidia/Nemotron-3-Diarization and nemotron-3.5-asr-streaming-0.6b-*.gguf from nvidia/nemotron-3.5-asr-streaming-0.6b. Both are governed by the OpenMDW License Agreement, version 1.1. The parakeet.cpp runtime is MIT-licensed.

Downloads last month
23,627
GGUF
Model size
0.6B params
Architecture
parakeet
Hardware compatibility
Log In to add your hardware

6-bit

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mudler/parakeet-cpp-gguf

Quantized
(3)
this model

Space using mudler/parakeet-cpp-gguf 1