Ekko v1 Tiny — offline Danish speech recognition

Leaderboard · Benchmark data · Evaluation harness

Ekko v1 Tiny — Offline Danish Speech-to-Text with Word Timestamps

Ekko v1 Tiny is a compact, 110M-parameter Danish speech recognition model with word timestamps. It is a FastConformer RNN-T fine-tuned from nvidia/parakeet-rnnt-110m-da-dk and runs locally. Its raw output is lowercase text; pair it with RyeAI/ekko-pnc for capitalization and punctuation.

Best fit: clear Danish speech from one speaker, delivered at a steady pace, such as dictation or reading aloud. The published results are stronger on read and formal speech than on spontaneous conversation. Dictation itself was not evaluated as a separate test set.

Benchmark

12.37% mean WER across 26,780 rows from five Danish test sets, evaluated with the leaderboard harness at commit 6257737. Scores use the published defaults: NFKC normalization, spoken-number folding, Danish number-word normalization, lowercase, punctuation removal, symmetric filler removal, and whitespace collapse. Lower is better.

Test set Speech type WER
CoRal conversation Spontaneous conversation 24.15%
CoRal read aloud Read speech 11.87%
Common Voice 17 Crowd-sourced read speech 8.44%
FLEURS da_dk Read news-style speech 9.75%
FTSpeech Parliamentary speech 7.62%
Equal-weight mean 12.37%

The NeMo checkpoint processed the benchmark at 795.8x realtime on one A100. The official run saved raw per-sample outputs, and an independent offline rescore reproduced 12.37% exactly.

Native GGUF quantization

The F32, F16, Q8_0, Q6_K, Q5_K, Q5_0, Q4_K, and Q4_0 files are scored on the same current 26,780-row matrix and normalization policy as the NeMo result above. All 40 variant/domain outputs completed and passed exact-row offline scoring.

Variant Size Mean WER CPU speed
F16 248.2 MiB 12.315% 74.4x
F32 430.6 MiB 12.318% 71.8x
Q8_0 162.7 MiB 12.323% 74.9x
Q6_K 141.8 MiB 12.325% 63.0x
Q5_K 129.8 MiB 12.337% 60.1x
Q5_0 128.5 MiB 12.370% 69.4x
Q4_K 118.4 MiB 12.442% 65.4x
Q4_0 117.1 MiB 12.467% 72.1x

Q5_0 remains the packaged default: it is 128.5 MiB, runs at 69.4x realtime, and is 0.047 mean-WER points behind Q8_0. Choose Q8_0 when the extra 34.2 MiB is acceptable and maximum quantized accuracy is preferred. F16 is the native reference tier.

The matrix uses leaderboard commit 6257737 and normalizer SHA-256 90fb9d8985f5919dffdb969a54cdd0f858f63d99fa6c0783921966a785f62c39. The saved filler-normalized score report has SHA-256 46fd3394b52a0a330ba3f8bb318417fdabe9ee7a1fa51358009299ebba661a31.

ONNX and browser export

The model is available for sherpa-onnx and was evaluated on the same 26,780-row matrix with mandatory non-silent peak normalization to 0.95. The browser downloads the compact int8 export.

Export Size Mean WER CPU speed
FP32 ONNX 450.6 MiB 12.465% 53.4x
INT8 ONNX 129.0 MiB 12.946% 49.2x

Both exports completed all five domains with zero runtime failures. INT8 is 71.4% smaller than FP32 and trades 0.481 mean-WER points for the smaller browser download. The filler-normalized ONNX score report has SHA-256 e26343000384349a721da11fe316136b92fb87ffa78cb19133b6cb7193831155.

Artifacts

The current NeMo, GGUF, and ONNX artifacts are pinned at Hub revision e37e698805c538e447c25062ab13fa862462ee9e and verified by SHA-256.

Clean-cache packaged download and inference tests passed on macOS arm64 and Linux x64 with native word timestamps. The Windows x64 artifact is mapped but has not been exercised against this release.

Format Path Measured mean WER
NeMo ekko-v1-tiny.nemo 12.37%
parakeet.cpp GGUF gguf-parakeet.cpp/ekko-v1-tiny-{f32,f16,q8_0,q6_k,q5_k,q5_0,q4_k,q4_0}.gguf 12.315-12.467%
sherpa-onnx FP32 onnx-sherpa/{encoder,decoder,joiner}.onnx 12.465%
sherpa-onnx INT8 onnx-sherpa/{encoder,decoder,joiner}.int8.onnx 12.946%

The browser loads onnx-sherpa/runtime-config.json with the INT8 files and checks its SHA-256 before using its audio and model settings.

Local Use

hf download RyeAI/ekko-v1-tiny \
  gguf-parakeet.cpp/ekko-v1-tiny-q5_0.gguf \
  --local-dir ./ekko-v1-tiny
parakeet-cli transcribe \
  --model ./ekko-v1-tiny/gguf-parakeet.cpp/ekko-v1-tiny-q5_0.gguf \
  --input audio.wav --timestamps

NeMo usage:

from nemo.collections.asr.models import EncDecRNNTBPEModel

model = EncDecRNNTBPEModel.from_pretrained("RyeAI/ekko-v1-tiny")
transcript = model.transcribe(["audio.wav"])[0]
print(transcript.text)

Training and limitations

Ekko v1 Tiny trained for 50,000 steps on an 879,083-row, 1,549.6-hour Danish manifest, including 454.1 hours of quality-gated pseudo-labelled audio.

  • Danish only.
  • Raw output has no punctuation or capitalization.
  • The 44-token vocabulary does not contain uppercase or punctuation tokens.
  • Punctuation and capitalization require a separate restoration model such as RyeAI/ekko-pnc.
  • Spontaneous conversation remains the hardest benchmark domain (24.15% WER).
  • Noisy audio and overlapping speakers were not evaluated separately.

License

The complete model and data terms are documented in MODEL_LICENSE.md and the bundled LICENSES/ files.

Downloads last month
30
GGUF
Model size
0.1B params
Architecture
parakeet
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RyeAI/ekko-v1-tiny

Finetuned
(2)
this model

Datasets used to train RyeAI/ekko-v1-tiny

Space using RyeAI/ekko-v1-tiny 1

Evaluation results