Automatic Speech Recognition
NeMo
ONNX
GGUF
Danish
speech-recognition
speech-to-text
danish
danish-asr
offline
local-inference
cpu-inference
sherpa-onnx
rnnt
word-timestamps
Eval Results (legacy)
Instructions to use RyeAI/ekko-v1-tiny with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use RyeAI/ekko-v1-tiny with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("RyeAI/ekko-v1-tiny") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from RyeAI/ekko-v1-tiny: direct link, hf CLI and curl.
- Browser
- Download file 7.87 kB
-
https://huggingface.co/RyeAI/ekko-v1-tiny/resolve/main/README.md
- Command line
-
hf download hf://RyeAI/ekko-v1-tiny/README.md
-
curl -L -o README.md https://huggingface.co/RyeAI/ekko-v1-tiny/resolve/main/README.md
7.87 kB
| language: | |
| - da | |
| license: other | |
| license_name: ekko-composite-model-terms | |
| license_link: https://huggingface.co/RyeAI/ekko-v1-tiny/blob/main/MODEL_LICENSE.md | |
| thumbnail: https://huggingface.co/RyeAI/ekko-v1-tiny/resolve/e37e698805c538e447c25062ab13fa862462ee9e/cover.png | |
| library_name: nemo | |
| pipeline_tag: automatic-speech-recognition | |
| base_model: nvidia/parakeet-rnnt-110m-da-dk | |
| base_model_relation: finetune | |
| datasets: | |
| - alexandrainst/nst-da | |
| - alexandrainst/ftspeech | |
| - alexandrainst/nota | |
| - CoRal-project/coral-v3 | |
| tags: | |
| - automatic-speech-recognition | |
| - speech-recognition | |
| - speech-to-text | |
| - danish | |
| - danish-asr | |
| - offline | |
| - local-inference | |
| - cpu-inference | |
| - nemo | |
| - gguf | |
| - onnx | |
| - sherpa-onnx | |
| - rnnt | |
| - word-timestamps | |
| metrics: | |
| - wer | |
| model-index: | |
| - name: Ekko v1 Tiny | |
| results: | |
| - task: | |
| type: automatic-speech-recognition | |
| dataset: | |
| name: CoRal conversation | |
| type: CoRal-project/coral-v3 | |
| config: conversation | |
| split: test | |
| metrics: | |
| - type: wer | |
| value: 24.15 | |
| - task: | |
| type: automatic-speech-recognition | |
| dataset: | |
| name: CoRal read aloud | |
| type: CoRal-project/coral-v3 | |
| config: read_aloud | |
| split: test | |
| metrics: | |
| - type: wer | |
| value: 11.87 | |
| - task: | |
| type: automatic-speech-recognition | |
| dataset: | |
| name: FLEURS Danish | |
| type: google/fleurs | |
| config: da_dk | |
| split: test | |
| metrics: | |
| - type: wer | |
| value: 9.75 | |
| - task: | |
| type: automatic-speech-recognition | |
| dataset: | |
| name: FTSpeech | |
| type: alexandrainst/ftspeech | |
| split: test_balanced | |
| metrics: | |
| - type: wer | |
| value: 7.62 | |
|  | |
| [Leaderboard](https://huggingface.co/spaces/RyeAI/danish-asr-leaderboard) · | |
| [Benchmark data](https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard) · | |
| [Evaluation harness](https://github.com/Rye-A1/danish-asr-leaderboard) | |
| # Ekko v1 Tiny — Offline Danish Speech-to-Text with Word Timestamps | |
| Ekko v1 Tiny is a compact, 110M-parameter Danish speech recognition model with | |
| word timestamps. It is a FastConformer RNN-T fine-tuned from | |
| `nvidia/parakeet-rnnt-110m-da-dk` and runs locally. Its raw output is lowercase | |
| text; pair it with `RyeAI/ekko-pnc` for capitalization and punctuation. | |
| **Best fit:** clear Danish speech from one speaker, delivered at a steady pace, | |
| such as dictation or reading aloud. The published results are stronger on read | |
| and formal speech than on spontaneous conversation. Dictation itself was not | |
| evaluated as a separate test set. | |
| ## Benchmark | |
| **12.37% mean WER** across 26,780 rows from five Danish test sets, evaluated | |
| with the leaderboard harness. Scores use the published defaults: NFKC | |
| normalization, spoken-number folding, Danish number-word normalization, | |
| lowercase, punctuation removal, symmetric filler removal, and whitespace | |
| collapse. Lower is better. | |
| | Test set | Speech type | WER | | |
| |---|---|---:| | |
| | CoRal conversation | Spontaneous conversation | 24.15% | | |
| | CoRal read aloud | Read speech | 11.87% | | |
| | Common Voice Danish (CV 25.0 test subset) | Crowd-sourced read speech | 8.44% | | |
| | FLEURS `da_dk` | Read news-style speech | 9.75% | | |
| | FTSpeech | Parliamentary speech | 7.62% | | |
| | **Equal-weight mean** | | **12.37%** | | |
| The NeMo checkpoint processed the benchmark at **795.8x realtime** on one A100. | |
| The official run saved raw per-sample outputs, and an independent offline | |
| rescore reproduced 12.37% exactly. | |
| ### Comparison with models below 800M parameters | |
|  | |
| The chart compares Tiny's Q5_0 release score with 12 open-weight models below | |
| 800M parameters in the [Danish ASR Leaderboard](https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard/tree/f7c930a30ac2f88e68421d88fc577c0cbafa8c31/results) | |
| as of 9 October 2026. Each score is an equal-weight mean WER across the same | |
| five Danish test sets (26,780 clips). Tiny scores 12.37%, 0.27 percentage | |
| points lower than Svale 600M at 12.64%. Rankings may change when new models | |
| or reruns are added. | |
| ### Native GGUF quantization | |
| The F32, F16, Q8_0, Q6_K, Q5_K, Q5_0, Q4_K, and Q4_0 files are | |
| scored on the same current 26,780-row matrix and normalization policy as the | |
| NeMo result above. All 40 variant/domain outputs completed and passed exact-row | |
| offline scoring. | |
| | Variant | Size | Mean WER | CPU speed | | |
| |---|---:|---:|---:| | |
| | **F16** | 248.2 MiB | **12.315%** | 74.4x | | |
| | F32 | 430.6 MiB | 12.318% | 71.8x | | |
| | **Q8_0** | 162.7 MiB | **12.323%** | **74.9x** | | |
| | Q6_K | 141.8 MiB | 12.325% | 63.0x | | |
| | Q5_K | 129.8 MiB | 12.337% | 60.1x | | |
| | **Q5_0** | **128.5 MiB** | **12.370%** | **69.4x** | | |
| | Q4_K | 118.4 MiB | 12.442% | 65.4x | | |
| | Q4_0 | 117.1 MiB | 12.467% | 72.1x | | |
| Q5_0 remains the packaged default: it is 128.5 MiB, runs at 69.4x realtime, | |
| and is 0.047 mean-WER points behind Q8_0. Choose Q8_0 when the extra 34.2 MiB | |
| is acceptable and maximum quantized accuracy is preferred. F16 is the native | |
| reference tier. | |
| ### ONNX and browser export | |
| The model is available for sherpa-onnx and was evaluated on the same 26,780-row | |
| matrix with mandatory non-silent peak normalization to 0.95. The browser | |
| downloads the compact int8 export. | |
| | Export | Size | Mean WER | CPU speed | | |
| |---|---:|---:|---:| | |
| | FP32 ONNX | 450.6 MiB | 12.465% | 53.4x | | |
| | **INT8 ONNX** | **129.0 MiB** | **12.946%** | **49.2x** | | |
| Both exports completed all five domains with zero runtime failures. INT8 is | |
| 71.4% smaller than FP32 and trades 0.481 mean-WER points for the smaller browser | |
| download. | |
| ## Artifacts | |
| The current NeMo, GGUF, and ONNX artifacts are pinned at Hub revision | |
| `e37e698805c538e447c25062ab13fa862462ee9e` and verified by SHA-256. | |
| Clean-cache packaged download and inference tests passed on macOS arm64 and | |
| Linux x64 with native word timestamps. The Windows x64 artifact is mapped but | |
| has not been exercised against this release. | |
| | Format | Path | Measured mean WER | | |
| |---|---|---:| | |
| | NeMo | `ekko-v1-tiny.nemo` | 12.37% | | |
| | parakeet.cpp GGUF | `gguf-parakeet.cpp/ekko-v1-tiny-{f32,f16,q8_0,q6_k,q5_k,q5_0,q4_k,q4_0}.gguf` | 12.315-12.467% | | |
| | sherpa-onnx FP32 | `onnx-sherpa/{encoder,decoder,joiner}.onnx` | 12.465% | | |
| | sherpa-onnx INT8 | `onnx-sherpa/{encoder,decoder,joiner}.int8.onnx` | 12.946% | | |
| The browser loads `onnx-sherpa/runtime-config.json` with the INT8 files and | |
| checks its SHA-256 before using its audio and model settings. | |
| ## Local Use | |
| For pinned downloads, word timestamps, and optional PnC, see the | |
| [Ekko STT quickstart](https://github.com/Rye-A1/ekko-danish-stt#run-locally). | |
| If you already have `parakeet-cli` from parakeet.cpp, you can also run the | |
| GGUF model directly: | |
| ```bash | |
| hf download RyeAI/ekko-v1-tiny \ | |
| gguf-parakeet.cpp/ekko-v1-tiny-q5_0.gguf \ | |
| --local-dir ./ekko-v1-tiny | |
| parakeet-cli transcribe \ | |
| --model ./ekko-v1-tiny/gguf-parakeet.cpp/ekko-v1-tiny-q5_0.gguf \ | |
| --input audio.wav --timestamps | |
| ``` | |
| NeMo usage: | |
| ```python | |
| from nemo.collections.asr.models import EncDecRNNTBPEModel | |
| model = EncDecRNNTBPEModel.from_pretrained("RyeAI/ekko-v1-tiny") | |
| transcript = model.transcribe(["audio.wav"])[0] | |
| print(transcript.text) | |
| ``` | |
| ## Training and limitations | |
| Ekko v1 Tiny trained for 50,000 steps on an 879,083-row, | |
| 1,549.6-hour Danish manifest, including 454.1 hours of quality-gated | |
| pseudo-labelled audio. | |
| - Danish only. | |
| - Raw output has no punctuation or capitalization. | |
| - The 44-token vocabulary does not contain uppercase or punctuation tokens. | |
| - Punctuation and capitalization require a separate restoration model such as | |
| `RyeAI/ekko-pnc`. | |
| - Spontaneous conversation remains the hardest benchmark domain (24.15% WER). | |
| - Noisy audio and overlapping speakers were not evaluated separately. | |
| ## License | |
| The complete model and data terms are documented in | |
| [`MODEL_LICENSE.md`](MODEL_LICENSE.md) and the bundled `LICENSES/` files. | |