Dolphin-Base-Burmese-ASR ( Burmese )

This model is a fine-tuned version of DataoceanAI/dolphin-base for Automatic Speech Recognition (ASR) in Burmese .

Model Overview

  • Architecture: Dolphin Base (E-Branchformer Encoder + Transformer Decoder)
  • Parameters: ~140M
  • Task: Automatic Speech Recognition (ASR)
  • Language: Burmese (my)
  • Sampling Rate: 16,000 Hz (Mono)
  • Tokenizer / Vocab: Dolphin Multilingual BPE (<my><MM> tokens)

Dataset Details

The model was trained on a standardized, high-quality Burmese speech corpus:

  • Total Duration: ~22 Hours of audio
  • Total Utterances: 24,560 WAV files (16 kHz, 16-bit PCM, Mono)
  • Total Speakers: 13 Synthetic Burmese Speakers
  • Speakers: แ€•แ€ฎแ€š แŠ แ€แ€ซแ€…แ€ฌ แŠ แ€žแ€ฎแ€›แ€ญ แŠ แ€’แ€ฎแ€• แŠ แ€กแ€€แ€นแ€แ€›แ€ฌ แŠ แ€žแ€’แ€นแ€’แ€ซ แŠ แ€žแ€› แŠ แ€แ€ฌแ€›แ€ฌ แŠ แ€€แ€แ€ญ แŠ แ€žแ€ฏแ€ แŠ แ€•แ€žแ€ฌแ€’ แŠ แ€‚แ€ฎแ€ แŠ แ€”แ€”แ€นแ€’
  • Training & Validation: 11 speakers (20,699 train utterances, 2,300 validation utterances)
  • Test Set (Held-Out): 2 unseen speakers ( แ€‚แ€ฎแ€ & แ€”แ€”แ€นแ€’ , 1,561 utterances )
  • Text Preprocessing: Standardized Myanmar Unicode (\u1000-\u109F) with unified word segmentation and punctuation removal. The held-out test split evaluates zero-shot acoustic generalization across completely unseen voices.

Training Configuration

Parameter Value
Effective Batch Size 64 (8 per device x 4 gradient accumulation x 2 GPUs)
Peak Learning Rate 5.0e-5
Learning Rate Scheduler Cosine
Warmup Steps 200
Total Steps 2,000 (~6 Epochs)
Mixed Precision FP16
Gradient Checkpointing Enabled
Augmentation None

Evaluation Results

Stage / Evaluation Step Train Loss Val Loss WER (%) CER (%) SER (%) DER (%) IER (%) chrF
Baseline (Untrained) 0 - - 92.13 17.17 100.00 47.23 0.58 78.82
Validation Set 250 5.1032 - 38.18 6.85 83.00 6.52 3.18 88.04
Validation Set 500 3.5427 - 31.67 4.95 74.00 5.15 4.09 92.20
Validation Set 750 2.8435 - 27.88 4.36 72.00 4.39 4.39 93.28
Validation Set 1000 2.4685 - 25.91 3.79 63.00 4.09 3.94 94.30
Validation Set 1250 2.4317 - 25.76 3.85 64.00 5.00 3.48 94.17
Validation Set 1500 2.1806 - 23.48 3.28 62.00 3.48 3.64 95.01
Validation Set 1750 2.0335 - 22.73 3.00 62.00 4.24 3.64 95.45
Validation Set 2000 1.9487 - 22.88 3.00 62.00 3.94 3.79 95.54
UNSEEN TEST (Final) Final - - 23.51 4.42 78.54 4.48 2.43 92.54

Usage

import dolphin

# 1. Load fine-tuned model
model = dolphin.load_model("base", model_dir="path_to_model", device="cuda")

# 2. Transcribe
result = dolphin.transcribe(model, "audio.wav", lang_sym="my", region_sym="MM")
print(result.text_nospecial)
Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for thantzinphyo/Dolphin-Base-Burmese-ASR

Finetuned
(2)
this model

Dataset used to train thantzinphyo/Dolphin-Base-Burmese-ASR

Collection including thantzinphyo/Dolphin-Base-Burmese-ASR

Evaluation results

  • Validation CER on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    3.000
  • Validation WER on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    22.880
  • Validation chrF on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    95.540
  • Test CER (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    4.420
  • Test WER (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    23.510
  • Test chrF (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    92.540