Instructions to use danish-foundation-models/edda-v0.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use danish-foundation-models/edda-v0.2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.2")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("danish-foundation-models/edda-v0.2") model = AutoModelForSpeechSeq2Seq.from_pretrained("danish-foundation-models/edda-v0.2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Edda v0.2 — Danish speech recognition
Edda is openai/whisper-large-v3-turbo (809 M parameters, 4-layer
decoder) fully fine-tuned on 2,600 hours of transcribed Danish speech. On the
open Danish ASR leaderboard harness it scores a mean WER of
8.79 over the five test sets, the best published result at the time of release, at 18.5x real time on a single AMD MI250X GCD.
| test set | WER | CER |
|---|---|---|
| CoRal-v3 conversation | 15.54 | 8.93 |
| CoRal-v3 read-aloud | 9.59 | 3.64 |
| Common Voice Danish (leaderboard set, 2,756 clips) | 5.71 | 1.85 |
| FLEURS da_dk | 7.34 | 2.84 |
| FTSpeech | 5.77 | 3.22 |
| mean | 8.79 | 4.10 |
Scores from the leaderboard's own harness (Rye-A1/danish-asr-leaderboard),
beam search with 5 beams, which is this model's default (see Decoding).
What changed since v0.1 (9.20 mean WER): v0.2 is trained
to use Whisper's previous-text prompt, which v0.1 was not (it looped when given an initial_prompt or when faster-whisper's
default condition_on_previous_text was on); the released weights average four checkpoints instead of one; and decoding
defaults to beam search.
Usage
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.2", device="cuda",
chunk_length_s=30) # clips longer than 30 s are transcribed in 30 s windows
print(asr("clip.wav")["text"])
# optional: previous text or domain vocabulary in Whisper's prompt slot (faster-whisper: initial_prompt / hotwords)
prompt_ids = asr.tokenizer.get_prompt_ids("Folketinget, finansloven, Dansk Industri", return_tensors="pt").to(asr.device)
print(asr("clip.wav", generate_kwargs={"prompt_ids": prompt_ids})["text"])
The generation config defaults to Danish transcription (language="da", task="transcribe") and to beam search with
5 beams, so neither needs to be passed. Recordings longer than 30 s are not truncated: with chunk_length_s=30 they are
transcribed in consecutive 30 s windows and joined, which is also how the leaderboard scores long clips. Any sample rate the
pipeline can read is accepted; audio is resampled to 16 kHz. Output is cased and punctuated in the style of the training
corpora: FTSpeech is lower-case without punctuation, the other corpora are cased and mostly punctuated.
Prompting. Edda v0.2 is trained to use Whisper's previous-text prompt. Half of the training clips were conditioned on
the same speaker's preceding utterance in the <|startofprev|> slot, so the decoder treats a prompt as context: the previous
segment's transcript helps, an unrelated prompt costs little, and neither derails it. This is the slot that initial_prompt
and hotwords fill in faster-whisper, and that its default condition_on_previous_text=True fills with the previous window's
output. Measured on the five leaderboard test sets with our harness, 5 beams:
| prompt | Edda v0.2 | Edda v0.1 |
|---|---|---|
| none | 8.76 | 9.15 |
the previous clip's own transcript (condition_on_previous_text=True) |
8.74 | 211 — 95 % of clips loop |
the same unrelated sentence on every clip (initial_prompt) |
8.83 | 126 — 30 % of clips loop |
v0.2's repetition-loop rate stays at 0.2–0.3 % of clips in all three settings. v0.1's prompted runs were decoded greedily; its collapse is not a decoding artefact but the untrained slot. All of this was measured through the transformers prompt path; a CTranslate2 conversion for faster-whisper has not been tested by us.
Decoding. generation_config.json sets num_beams=5, the setting every number on this card was measured with; the
transformers speech-recognition pipeline and faster-whisper use 5 beams by default anyway, but direct model.generate() calls
and older transformers releases would otherwise decode greedily. Greedy decoding is about 0.3 WER worse and 1.3–1.7× faster.
Training data
Six public corpora, training splits only, gold transcripts only (no pseudo-labels), 1.62 M clips / 2,605 h after filtering:
| corpus | source | clips | hours | share of training steps |
|---|---|---|---|---|
| FTSpeech (parliament) | alexandrainst/ftspeech train |
983,989 | 1,690 | 43 % |
| CoRal-v3 read-aloud | CoRal-project/coral-v3 read_aloud/train |
299,253 | 521 | 24 % |
| NST Danish | alexandrainst/nst-da train¹ |
182,575 | 239 | 16 % |
| CoRal-v3 conversation | CoRal-project/coral-v3 conversation/train |
147,184 | 144 | 12 % |
| FLEURS da_dk | google/fleurs train |
2,463 | 7.5 | 3 % |
| Common Voice 17 Danish | mozilla-foundation/common_voice_17_0 train |
3,484 | 4.1 | 2 % |
¹ NST's own test split was held out even though it is not a benchmark set.
Prompt conditioning. For 50 % of training clips the transcript of the same speaker's preceding utterance (manifest order,
which is recording order for CoRal and FTSpeech) is placed in Whisper's <|startofprev|> slot before the task tokens, truncated
to its last 223 tokens; the loss is not computed on the prompt. The other half of the clips, and clips with no preceding
utterance by the same speaker, are seen without a prompt, so the model is equally at home with an empty slot.
Training recipe
Full fine-tune of encoder and decoder from openai/whisper-large-v3-turbo, unmodified architecture (learned absolute
positions, fixed 30 s windows). One long constant-learning-rate run with four exits, each cooled down separately, then averaged.
| base run | 80,000 steps × 256 clips at a constant 1e-5 after a 500-step warmup; full training state saved at steps 20,000, 40,000, 60,000 and 80,000 (8 nodes × 8 MI250X GCDs × 4 clips; 20 h) |
| cooldown branches | from each saved state, 10,000 further steps with the learning rate decayed linearly to 1e-6 (4 nodes × 8 GCDs × 8 clips, 3 h 50 m each) |
| released weights | the four branch endpoints (EMA 0.9998 weights at steps 30,000, 50,000, 70,000 and 90,000) averaged with equal weights in fp32, stored in fp16 |
| optimizer | AdamW, β (0.9, 0.98), weight decay 0.01, gradient clipping 1.0, bf16 autocast |
| loss | cross-entropy with label smoothing 0.05 on `< |
| SpecAugment | time masks p 0.05 × 10 frames, feature masks p 0.05 × 10 bins |
| audio augmentation (per clip) | speed 0.9–1.1 (p 0.5); additive coloured noise at 0–20 dB SNR (p 0.6); 0.5–2 s trailing silence (p 0.25); 2.5 % of clips replaced by low-level noise with an empty transcript |
The single branch endpoints score 9.21, 9.02, 8.98 and 9.22 mean WER with greedy decoding on our harness; averaging the four and decoding with 5 beams gives the released 8.79.
Limitations
- Transcription convention follows the corpora. FTSpeech references are the edited parliamentary record, so on parliamentary speech the model tends to omit restarts and repetitions; CoRal references are verbatim.
- Long audio is handled by fixed 30 s windows; on very long silences or music the decoder can occasionally repeat a phrase (0.2–0.3 % of test clips). Split at pauses for best results.
- Domains not represented in training (children, strong non-native accents, telephone-band audio, singing) are untested.
- Evaluation references have known noise: a small fraction of CoRal read-aloud test prompts were paraphrased by the speaker, and CoRal conversation references systematically omit the copula er; both count against every model equally.
License and attribution
The model weights are released under the Apache License 2.0 by the Alexandra Institute, which is also the licensor of the
CoRal-v3 dataset. The base model openai/whisper-large-v3-turbo
is MIT-licensed; the training corpora carry their own licenses — see the respective dataset cards. Trained by the Alexandra
Institute within the CoRal project and released by Danish Foundation Models.
- Downloads last month
- 157
Model tree for danish-foundation-models/edda-v0.2
Datasets used to train danish-foundation-models/edda-v0.2
CoRal-project/coral-v3
alexandrainst/ftspeech
Spaces using danish-foundation-models/edda-v0.2 2
Evaluation results
- WER on CoRal-v3 conversation (test)test set self-reported15.540
- CER on CoRal-v3 conversation (test)test set self-reported8.930
- WER on CoRal-v3 read-aloud (test)test set self-reported9.590
- CER on CoRal-v3 read-aloud (test)test set self-reported3.640
- WER on Common Voice Danish (leaderboard settest set self-reported5.710
- CER on Common Voice Danish (leaderboard settest set self-reported1.850
- WER on FLEURS da_dk (test)test set self-reported7.340
- CER on FLEURS da_dk (test)test set self-reported2.840