ekko-v1-tiny / README.md
emilstabil's picture
Date comparison to current leaderboard results
950d88b verified
|
Raw History Blame Contribute Delete
7.87 kB
metadata
language:
  - da
license: other
license_name: ekko-composite-model-terms
license_link: https://huggingface.co/RyeAI/ekko-v1-tiny/blob/main/MODEL_LICENSE.md
thumbnail: >-
  https://huggingface.co/RyeAI/ekko-v1-tiny/resolve/e37e698805c538e447c25062ab13fa862462ee9e/cover.png
library_name: nemo
pipeline_tag: automatic-speech-recognition
base_model: nvidia/parakeet-rnnt-110m-da-dk
base_model_relation: finetune
datasets:
  - alexandrainst/nst-da
  - alexandrainst/ftspeech
  - alexandrainst/nota
  - CoRal-project/coral-v3
tags:
  - automatic-speech-recognition
  - speech-recognition
  - speech-to-text
  - danish
  - danish-asr
  - offline
  - local-inference
  - cpu-inference
  - nemo
  - gguf
  - onnx
  - sherpa-onnx
  - rnnt
  - word-timestamps
metrics:
  - wer
model-index:
  - name: Ekko v1 Tiny
    results:
      - task:
          type: automatic-speech-recognition
        dataset:
          name: CoRal conversation
          type: CoRal-project/coral-v3
          config: conversation
          split: test
        metrics:
          - type: wer
            value: 24.15
      - task:
          type: automatic-speech-recognition
        dataset:
          name: CoRal read aloud
          type: CoRal-project/coral-v3
          config: read_aloud
          split: test
        metrics:
          - type: wer
            value: 11.87
      - task:
          type: automatic-speech-recognition
        dataset:
          name: FLEURS Danish
          type: google/fleurs
          config: da_dk
          split: test
        metrics:
          - type: wer
            value: 9.75
      - task:
          type: automatic-speech-recognition
        dataset:
          name: FTSpeech
          type: alexandrainst/ftspeech
          split: test_balanced
        metrics:
          - type: wer
            value: 7.62

Ekko v1 Tiny — offline Danish speech recognition

Leaderboard · Benchmark data · Evaluation harness

Ekko v1 Tiny — Offline Danish Speech-to-Text with Word Timestamps

Ekko v1 Tiny is a compact, 110M-parameter Danish speech recognition model with word timestamps. It is a FastConformer RNN-T fine-tuned from nvidia/parakeet-rnnt-110m-da-dk and runs locally. Its raw output is lowercase text; pair it with RyeAI/ekko-pnc for capitalization and punctuation.

Best fit: clear Danish speech from one speaker, delivered at a steady pace, such as dictation or reading aloud. The published results are stronger on read and formal speech than on spontaneous conversation. Dictation itself was not evaluated as a separate test set.

Benchmark

12.37% mean WER across 26,780 rows from five Danish test sets, evaluated with the leaderboard harness. Scores use the published defaults: NFKC normalization, spoken-number folding, Danish number-word normalization, lowercase, punctuation removal, symmetric filler removal, and whitespace collapse. Lower is better.

Test set Speech type WER
CoRal conversation Spontaneous conversation 24.15%
CoRal read aloud Read speech 11.87%
Common Voice Danish (CV 25.0 test subset) Crowd-sourced read speech 8.44%
FLEURS da_dk Read news-style speech 9.75%
FTSpeech Parliamentary speech 7.62%
Equal-weight mean 12.37%

The NeMo checkpoint processed the benchmark at 795.8x realtime on one A100. The official run saved raw per-sample outputs, and an independent offline rescore reproduced 12.37% exactly.

Comparison with models below 800M parameters

Mean WER of Ekko v1 Tiny and all twelve open-weight leaderboard models below 800M parameters

The chart compares Tiny's Q5_0 release score with 12 open-weight models below 800M parameters in the Danish ASR Leaderboard as of 9 October 2026. Each score is an equal-weight mean WER across the same five Danish test sets (26,780 clips). Tiny scores 12.37%, 0.27 percentage points lower than Svale 600M at 12.64%. Rankings may change when new models or reruns are added.

Native GGUF quantization

The F32, F16, Q8_0, Q6_K, Q5_K, Q5_0, Q4_K, and Q4_0 files are scored on the same current 26,780-row matrix and normalization policy as the NeMo result above. All 40 variant/domain outputs completed and passed exact-row offline scoring.

Variant Size Mean WER CPU speed
F16 248.2 MiB 12.315% 74.4x
F32 430.6 MiB 12.318% 71.8x
Q8_0 162.7 MiB 12.323% 74.9x
Q6_K 141.8 MiB 12.325% 63.0x
Q5_K 129.8 MiB 12.337% 60.1x
Q5_0 128.5 MiB 12.370% 69.4x
Q4_K 118.4 MiB 12.442% 65.4x
Q4_0 117.1 MiB 12.467% 72.1x

Q5_0 remains the packaged default: it is 128.5 MiB, runs at 69.4x realtime, and is 0.047 mean-WER points behind Q8_0. Choose Q8_0 when the extra 34.2 MiB is acceptable and maximum quantized accuracy is preferred. F16 is the native reference tier.

ONNX and browser export

The model is available for sherpa-onnx and was evaluated on the same 26,780-row matrix with mandatory non-silent peak normalization to 0.95. The browser downloads the compact int8 export.

Export Size Mean WER CPU speed
FP32 ONNX 450.6 MiB 12.465% 53.4x
INT8 ONNX 129.0 MiB 12.946% 49.2x

Both exports completed all five domains with zero runtime failures. INT8 is 71.4% smaller than FP32 and trades 0.481 mean-WER points for the smaller browser download.

Artifacts

The current NeMo, GGUF, and ONNX artifacts are pinned at Hub revision e37e698805c538e447c25062ab13fa862462ee9e and verified by SHA-256.

Clean-cache packaged download and inference tests passed on macOS arm64 and Linux x64 with native word timestamps. The Windows x64 artifact is mapped but has not been exercised against this release.

Format Path Measured mean WER
NeMo ekko-v1-tiny.nemo 12.37%
parakeet.cpp GGUF gguf-parakeet.cpp/ekko-v1-tiny-{f32,f16,q8_0,q6_k,q5_k,q5_0,q4_k,q4_0}.gguf 12.315-12.467%
sherpa-onnx FP32 onnx-sherpa/{encoder,decoder,joiner}.onnx 12.465%
sherpa-onnx INT8 onnx-sherpa/{encoder,decoder,joiner}.int8.onnx 12.946%

The browser loads onnx-sherpa/runtime-config.json with the INT8 files and checks its SHA-256 before using its audio and model settings.

Local Use

For pinned downloads, word timestamps, and optional PnC, see the Ekko STT quickstart. If you already have parakeet-cli from parakeet.cpp, you can also run the GGUF model directly:

hf download RyeAI/ekko-v1-tiny \
  gguf-parakeet.cpp/ekko-v1-tiny-q5_0.gguf \
  --local-dir ./ekko-v1-tiny
parakeet-cli transcribe \
  --model ./ekko-v1-tiny/gguf-parakeet.cpp/ekko-v1-tiny-q5_0.gguf \
  --input audio.wav --timestamps

NeMo usage:

from nemo.collections.asr.models import EncDecRNNTBPEModel

model = EncDecRNNTBPEModel.from_pretrained("RyeAI/ekko-v1-tiny")
transcript = model.transcribe(["audio.wav"])[0]
print(transcript.text)

Training and limitations

Ekko v1 Tiny trained for 50,000 steps on an 879,083-row, 1,549.6-hour Danish manifest, including 454.1 hours of quality-gated pseudo-labelled audio.

  • Danish only.
  • Raw output has no punctuation or capitalization.
  • The 44-token vocabulary does not contain uppercase or punctuation tokens.
  • Punctuation and capitalization require a separate restoration model such as RyeAI/ekko-pnc.
  • Spontaneous conversation remains the hardest benchmark domain (24.15% WER).
  • Noisy audio and overlapping speakers were not evaluated separately.

License

The complete model and data terms are documented in MODEL_LICENSE.md and the bundled LICENSES/ files.