Instructions to use RyeAI/ekko-v1-tiny with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use RyeAI/ekko-v1-tiny with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("RyeAI/ekko-v1-tiny") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Leaderboard · Benchmark data · Evaluation harness
Ekko v1 Tiny — Offline Danish Speech-to-Text with Word Timestamps
Ekko v1 Tiny is a compact, 110M-parameter Danish speech recognition model with
word timestamps. It is a FastConformer RNN-T fine-tuned from
nvidia/parakeet-rnnt-110m-da-dk and runs locally. Its raw output is lowercase
text; pair it with RyeAI/ekko-pnc for capitalization and punctuation.
Best fit: clear Danish speech from one speaker, delivered at a steady pace, such as dictation or reading aloud. The published results are stronger on read and formal speech than on spontaneous conversation. Dictation itself was not evaluated as a separate test set.
Benchmark
12.37% mean WER across 26,780 rows from five Danish test sets, evaluated
with the leaderboard harness at commit 6257737. Scores use
the published defaults: NFKC normalization, spoken-number folding, Danish
number-word normalization, lowercase, punctuation removal, symmetric filler
removal, and whitespace collapse. Lower is better.
| Test set | Speech type | WER |
|---|---|---|
| CoRal conversation | Spontaneous conversation | 24.15% |
| CoRal read aloud | Read speech | 11.87% |
| Common Voice 17 | Crowd-sourced read speech | 8.44% |
FLEURS da_dk |
Read news-style speech | 9.75% |
| FTSpeech | Parliamentary speech | 7.62% |
| Equal-weight mean | 12.37% |
The NeMo checkpoint processed the benchmark at 795.8x realtime on one A100. The official run saved raw per-sample outputs, and an independent offline rescore reproduced 12.37% exactly.
Native GGUF quantization
The F32, F16, Q8_0, Q6_K, Q5_K, Q5_0, Q4_K, and Q4_0 files are scored on the same current 26,780-row matrix and normalization policy as the NeMo result above. All 40 variant/domain outputs completed and passed exact-row offline scoring.
| Variant | Size | Mean WER | CPU speed |
|---|---|---|---|
| F16 | 248.2 MiB | 12.315% | 74.4x |
| F32 | 430.6 MiB | 12.318% | 71.8x |
| Q8_0 | 162.7 MiB | 12.323% | 74.9x |
| Q6_K | 141.8 MiB | 12.325% | 63.0x |
| Q5_K | 129.8 MiB | 12.337% | 60.1x |
| Q5_0 | 128.5 MiB | 12.370% | 69.4x |
| Q4_K | 118.4 MiB | 12.442% | 65.4x |
| Q4_0 | 117.1 MiB | 12.467% | 72.1x |
Q5_0 remains the packaged default: it is 128.5 MiB, runs at 69.4x realtime, and is 0.047 mean-WER points behind Q8_0. Choose Q8_0 when the extra 34.2 MiB is acceptable and maximum quantized accuracy is preferred. F16 is the native reference tier.
The matrix uses leaderboard commit 6257737 and normalizer SHA-256
90fb9d8985f5919dffdb969a54cdd0f858f63d99fa6c0783921966a785f62c39.
The saved filler-normalized score report has SHA-256
46fd3394b52a0a330ba3f8bb318417fdabe9ee7a1fa51358009299ebba661a31.
ONNX and browser export
The model is available for sherpa-onnx and was evaluated on the same 26,780-row matrix with mandatory non-silent peak normalization to 0.95. The browser downloads the compact int8 export.
| Export | Size | Mean WER | CPU speed |
|---|---|---|---|
| FP32 ONNX | 450.6 MiB | 12.465% | 53.4x |
| INT8 ONNX | 129.0 MiB | 12.946% | 49.2x |
Both exports completed all five domains with zero runtime failures. INT8 is
71.4% smaller than FP32 and trades 0.481 mean-WER points for the smaller browser
download. The filler-normalized ONNX score report has SHA-256
e26343000384349a721da11fe316136b92fb87ffa78cb19133b6cb7193831155.
Artifacts
The current NeMo, GGUF, and ONNX artifacts are pinned at Hub revision
e37e698805c538e447c25062ab13fa862462ee9e and verified by SHA-256.
Clean-cache packaged download and inference tests passed on macOS arm64 and Linux x64 with native word timestamps. The Windows x64 artifact is mapped but has not been exercised against this release.
| Format | Path | Measured mean WER |
|---|---|---|
| NeMo | ekko-v1-tiny.nemo |
12.37% |
| parakeet.cpp GGUF | gguf-parakeet.cpp/ekko-v1-tiny-{f32,f16,q8_0,q6_k,q5_k,q5_0,q4_k,q4_0}.gguf |
12.315-12.467% |
| sherpa-onnx FP32 | onnx-sherpa/{encoder,decoder,joiner}.onnx |
12.465% |
| sherpa-onnx INT8 | onnx-sherpa/{encoder,decoder,joiner}.int8.onnx |
12.946% |
The browser loads onnx-sherpa/runtime-config.json with the INT8 files and
checks its SHA-256 before using its audio and model settings.
Local Use
hf download RyeAI/ekko-v1-tiny \
gguf-parakeet.cpp/ekko-v1-tiny-q5_0.gguf \
--local-dir ./ekko-v1-tiny
parakeet-cli transcribe \
--model ./ekko-v1-tiny/gguf-parakeet.cpp/ekko-v1-tiny-q5_0.gguf \
--input audio.wav --timestamps
NeMo usage:
from nemo.collections.asr.models import EncDecRNNTBPEModel
model = EncDecRNNTBPEModel.from_pretrained("RyeAI/ekko-v1-tiny")
transcript = model.transcribe(["audio.wav"])[0]
print(transcript.text)
Training and limitations
Ekko v1 Tiny trained for 50,000 steps on an 879,083-row, 1,549.6-hour Danish manifest, including 454.1 hours of quality-gated pseudo-labelled audio.
- Danish only.
- Raw output has no punctuation or capitalization.
- The 44-token vocabulary does not contain uppercase or punctuation tokens.
- Punctuation and capitalization require a separate restoration model such as
RyeAI/ekko-pnc. - Spontaneous conversation remains the hardest benchmark domain (24.15% WER).
- Noisy audio and overlapping speakers were not evaluated separately.
License
The complete model and data terms are documented in
MODEL_LICENSE.md and the bundled LICENSES/ files.
- Downloads last month
- 30
4-bit
5-bit
6-bit
8-bit
16-bit
32-bit
Model tree for RyeAI/ekko-v1-tiny
Base model
nvidia/parakeet-rnnt-110m-da-dkDatasets used to train RyeAI/ekko-v1-tiny
alexandrainst/ftspeech
alexandrainst/nst-da
Space using RyeAI/ekko-v1-tiny 1
Evaluation results
- wer on CoRal conversationtest set self-reported24.150
- wer on CoRal read aloudtest set self-reported11.870
- wer on Common Voice 17 Danishtest set self-reported8.440
- wer on FLEURS Danishtest set self-reported9.750
- wer on FTSpeechself-reported7.620
