Dolphin-Small-Burmese-ASR ( Burmese )

This model is a fine-tuned version of DataoceanAI/dolphin-small for Automatic Speech Recognition (ASR) in Burmese.

Model Overview

  • Architecture: Dolphin Small (E-Branchformer Encoder + Transformer Decoder, 372.3M params)
  • Task: Automatic Speech Recognition (ASR)
  • Language: Burmese (my)
  • Sampling Rate: 16,000 Hz (Mono)
  • Tokenizer / Vocab: Dolphin Multilingual BPE (<my><MM> tokens)

Dataset Details

The model was trained on a standardized, high-quality Burmese speech corpus:

  • Total Duration: ~22 Hours of audio
  • Total Utterances: 24,560 WAV files (16 kHz, 16-bit PCM, Mono)
  • Total Speakers: 13 Synthetic Burmese Speakers
  • Speakers: แ€•แ€ฎแ€š แŠ แ€แ€ซแ€…แ€ฌ แŠ แ€žแ€ฎแ€›แ€ญ แŠ แ€’แ€ฎแ€• แŠ แ€กแ€€แ€นแ€แ€›แ€ฌ แŠ แ€žแ€’แ€นแ€’แ€ซ แŠ แ€žแ€› แŠ แ€แ€ฌแ€›แ€ฌ แŠ แ€€แ€แ€ญ แŠ แ€žแ€ฏแ€ แŠ แ€•แ€žแ€ฌแ€’ แŠ แ€‚แ€ฎแ€ แŠ แ€”แ€”แ€นแ€’
  • Training & Validation: 11 speakers (20,699 train utterances, 2,300 validation utterances)
  • Test Set (Held-Out): 2 unseen speakers ( แ€‚แ€ฎแ€ & แ€”แ€”แ€นแ€’ , 1,561 utterances )
  • Text Preprocessing: Standardized Myanmar Unicode (\u1000-\u109F) with unified word segmentation and punctuation removal. The held-out test split evaluates zero-shot acoustic generalization across completely unseen voices.

Training Configuration

Parameter Value
Effective Batch Size 64 (8 per device x 4 gradient accumulation x 2 GPUs)
Peak Learning Rate 5.0e-5
Learning Rate Scheduler Cosine
Warmup Steps 200
Total Steps 2,000 (~6 Epochs)
Mixed Precision FP16
Gradient Checkpointing Enabled
Augmentation None

Evaluation Results

Stage / Evaluation Step Train Loss Val Loss WER (%) CER (%) SER (%) DER (%) IER (%) chrF
Baseline (Untrained) 0 - - 95.34 12.85 100.00 53.94 0.00 88.48
Validation Set 250 3.6786 - 31.67 4.36 75.00 6.82 3.18 93.15
Validation Set 500 2.6207 - 25.91 3.77 68.00 4.70 3.48 94.06
Validation Set 750 2.4539 - 27.58 3.92 67.00 4.70 4.09 94.00
Validation Set 1000 1.9528 - 23.64 3.44 62.00 5.15 3.03 95.35
Validation Set 1250 2.0450 - 24.85 3.48 62.00 4.85 3.03 94.19
Validation Set 1500 1.7359 - 21.36 3.04 58.00 3.33 2.73 95.24
Validation Set 1750 1.7104 - 20.45 2.74 56.00 3.48 2.88 95.82
Validation Set 2000 1.6303 - 20.15 2.87 55.00 3.33 2.42 95.64
UNSEEN TEST (Final) Final - - 19.39 3.38 71.24 3.93 2.28 94.57

Usage

import dolphin

# 1. Load fine-tuned model
model = dolphin.load_model("small", model_dir="path_to_model", device="cuda")

# 2. Transcribe
result = dolphin.transcribe(model, "audio.wav", lang_sym="my", region_sym="MM")
print(result.text_nospecial)
Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for thantzinphyo/Dolphin-Small-Burmese-ASR

Finetuned
(2)
this model

Dataset used to train thantzinphyo/Dolphin-Small-Burmese-ASR

Evaluation results

  • Validation CER on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    2.870
  • Validation WER on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    20.150
  • Validation chrF on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    95.640
  • Test CER (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    3.380
  • Test WER (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    19.390
  • Test chrF (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    94.570