Edda v0.2 — Danish speech recognition

Edda is openai/whisper-large-v3-turbo (809 M parameters, 4-layer decoder) fully fine-tuned on 2,600 hours of transcribed Danish speech. On the open Danish ASR leaderboard harness it scores a mean WER of 8.79 over the five test sets, the best published result at the time of release, at 18.5x real time on a single AMD MI250X GCD.

test set WER CER
CoRal-v3 conversation 15.54 8.93
CoRal-v3 read-aloud 9.59 3.64
Common Voice Danish (leaderboard set, 2,756 clips) 5.71 1.85
FLEURS da_dk 7.34 2.84
FTSpeech 5.77 3.22
mean 8.79 4.10

Scores from the leaderboard's own harness (Rye-A1/danish-asr-leaderboard), beam search with 5 beams, which is this model's default (see Decoding).

What changed since v0.1 (9.20 mean WER): v0.2 is trained to use Whisper's previous-text prompt, which v0.1 was not (it looped when given an initial_prompt or when faster-whisper's default condition_on_previous_text was on); the released weights average four checkpoints instead of one; and decoding defaults to beam search.

Usage

from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.2", device="cuda",
               chunk_length_s=30)                         # clips longer than 30 s are transcribed in 30 s windows
print(asr("clip.wav")["text"])

# optional: previous text or domain vocabulary in Whisper's prompt slot (faster-whisper: initial_prompt / hotwords)
prompt_ids = asr.tokenizer.get_prompt_ids("Folketinget, finansloven, Dansk Industri", return_tensors="pt").to(asr.device)
print(asr("clip.wav", generate_kwargs={"prompt_ids": prompt_ids})["text"])

The generation config defaults to Danish transcription (language="da", task="transcribe") and to beam search with 5 beams, so neither needs to be passed. Recordings longer than 30 s are not truncated: with chunk_length_s=30 they are transcribed in consecutive 30 s windows and joined, which is also how the leaderboard scores long clips. Any sample rate the pipeline can read is accepted; audio is resampled to 16 kHz. Output is cased and punctuated in the style of the training corpora: FTSpeech is lower-case without punctuation, the other corpora are cased and mostly punctuated.

Prompting. Edda v0.2 is trained to use Whisper's previous-text prompt. Half of the training clips were conditioned on the same speaker's preceding utterance in the <|startofprev|> slot, so the decoder treats a prompt as context: the previous segment's transcript helps, an unrelated prompt costs little, and neither derails it. This is the slot that initial_prompt and hotwords fill in faster-whisper, and that its default condition_on_previous_text=True fills with the previous window's output. Measured on the five leaderboard test sets with our harness, 5 beams:

prompt Edda v0.2 Edda v0.1
none 8.76 9.15
the previous clip's own transcript (condition_on_previous_text=True) 8.74 211 — 95 % of clips loop
the same unrelated sentence on every clip (initial_prompt) 8.83 126 — 30 % of clips loop

v0.2's repetition-loop rate stays at 0.2–0.3 % of clips in all three settings. v0.1's prompted runs were decoded greedily; its collapse is not a decoding artefact but the untrained slot. All of this was measured through the transformers prompt path; a CTranslate2 conversion for faster-whisper has not been tested by us.

Decoding. generation_config.json sets num_beams=5, the setting every number on this card was measured with; the transformers speech-recognition pipeline and faster-whisper use 5 beams by default anyway, but direct model.generate() calls and older transformers releases would otherwise decode greedily. Greedy decoding is about 0.3 WER worse and 1.3–1.7× faster.

Training data

Six public corpora, training splits only, gold transcripts only (no pseudo-labels), 1.62 M clips / 2,605 h after filtering:

corpus source clips hours share of training steps
FTSpeech (parliament) alexandrainst/ftspeech train 983,989 1,690 43 %
CoRal-v3 read-aloud CoRal-project/coral-v3 read_aloud/train 299,253 521 24 %
NST Danish alexandrainst/nst-da train¹ 182,575 239 16 %
CoRal-v3 conversation CoRal-project/coral-v3 conversation/train 147,184 144 12 %
FLEURS da_dk google/fleurs train 2,463 7.5 3 %
Common Voice 17 Danish mozilla-foundation/common_voice_17_0 train 3,484 4.1 2 %

¹ NST's own test split was held out even though it is not a benchmark set.

Prompt conditioning. For 50 % of training clips the transcript of the same speaker's preceding utterance (manifest order, which is recording order for CoRal and FTSpeech) is placed in Whisper's <|startofprev|> slot before the task tokens, truncated to its last 223 tokens; the loss is not computed on the prompt. The other half of the clips, and clips with no preceding utterance by the same speaker, are seen without a prompt, so the model is equally at home with an empty slot.

Training recipe

Full fine-tune of encoder and decoder from openai/whisper-large-v3-turbo, unmodified architecture (learned absolute positions, fixed 30 s windows). One long constant-learning-rate run with four exits, each cooled down separately, then averaged.

base run 80,000 steps × 256 clips at a constant 1e-5 after a 500-step warmup; full training state saved at steps 20,000, 40,000, 60,000 and 80,000 (8 nodes × 8 MI250X GCDs × 4 clips; 20 h)
cooldown branches from each saved state, 10,000 further steps with the learning rate decayed linearly to 1e-6 (4 nodes × 8 GCDs × 8 clips, 3 h 50 m each)
released weights the four branch endpoints (EMA 0.9998 weights at steps 30,000, 50,000, 70,000 and 90,000) averaged with equal weights in fp32, stored in fp16
optimizer AdamW, β (0.9, 0.98), weight decay 0.01, gradient clipping 1.0, bf16 autocast
loss cross-entropy with label smoothing 0.05 on `<
SpecAugment time masks p 0.05 × 10 frames, feature masks p 0.05 × 10 bins
audio augmentation (per clip) speed 0.9–1.1 (p 0.5); additive coloured noise at 0–20 dB SNR (p 0.6); 0.5–2 s trailing silence (p 0.25); 2.5 % of clips replaced by low-level noise with an empty transcript

The single branch endpoints score 9.21, 9.02, 8.98 and 9.22 mean WER with greedy decoding on our harness; averaging the four and decoding with 5 beams gives the released 8.79.

Limitations

  • Transcription convention follows the corpora. FTSpeech references are the edited parliamentary record, so on parliamentary speech the model tends to omit restarts and repetitions; CoRal references are verbatim.
  • Long audio is handled by fixed 30 s windows; on very long silences or music the decoder can occasionally repeat a phrase (0.2–0.3 % of test clips). Split at pauses for best results.
  • Domains not represented in training (children, strong non-native accents, telephone-band audio, singing) are untested.
  • Evaluation references have known noise: a small fraction of CoRal read-aloud test prompts were paraphrased by the speaker, and CoRal conversation references systematically omit the copula er; both count against every model equally.

License and attribution

The model weights are released under the Apache License 2.0 by the Alexandra Institute, which is also the licensor of the CoRal-v3 dataset. The base model openai/whisper-large-v3-turbo is MIT-licensed; the training corpora carry their own licenses — see the respective dataset cards. Trained by the Alexandra Institute within the CoRal project and released by Danish Foundation Models.

Downloads last month
157
Safetensors
Model size
0.8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for danish-foundation-models/edda-v0.2

Finetuned
(644)
this model
Finetunes
2 models

Datasets used to train danish-foundation-models/edda-v0.2

Spaces using danish-foundation-models/edda-v0.2 2

Evaluation results