--- license: cc0-1.0 language: - my tags: - text-to-speech - tts - burmese - myanmar - stabletts - from-scratch pipeline_tag: text-to-speech --- # MyanmarTTS A from-scratch Burmese (Myanmar) text-to-speech model. 31M parameters, ~63 MB (fp16). - **Language**: Burmese (မြန်မာဘာသာ) - **Training data**: ~1.34M samples (news + real-world audio) - **Training steps**: 81,000 - **Training hardware**: A100-80GB (Colab Pro+), ~13 hours, ~88 compute units - **Inference hardware**: Free Colab T4, any modern GPU, or CPU (slower) - **Architecture**: StableTTS (DiT + flow matching) + Vocos vocoder - **License**: **CC0 1.0** (public domain, no attribution required) --- ## Sample Output All samples generated with default settings (`euler` solver, 12 steps). **Sample 0** **Text:** "မြန်မာလူမျိုးများဟာ အလွန် ယဥ်ကျေးသိမ်မွေ့ပြီး ဧည့်သည်များကို ပျူပျူငှာငှာ လှိုက်လှိုက်လှဲလှဲနဲ့ ကြိုဆိုကြပါတယ်" **Sample 1** **Text:** "မင်္ဂလာပါရှင် ကျွန်မကတော့ မြန်မာလူမျိုး ကရင်တိုင်းရင်းသူ အမျိုးသမီးလေး တစ်ဦး ဖြစ်ပါတယ်" **Sample 2** **Text:** "ဒီနေ့ ကျွန်မတို့ရဲ့ တီတီအက်စ် စနစ်သစ်လေး မော်ဒယ်အသစ်လေးတစ်ခုကို အောင်မြင်စွာ လေ့ကျင့် သင်ကြားနိုင်ခဲ့ပါတယ်" **Sample 3** **Text:** "လူသားတိုင်း လူသားတိုင်း ကိုယ်စိတ်နှစ်ဖြာ ကျန်းမာရွှင်လန်းပြီး စီးပွားလာဘ်လာဘတွေ ဒီရေအလား ကြီးပွား တိုးတက်နိုင်ကြပါစေ" **Sample 4** **Text:** "ဒီအသံထုတ်စနစ်လေးကို အသုံးပြုသူတိုင်း ကျန်းမာချမ်းသာပြီး လိုရာဆန္ဒတွေ တလုံးတဝတည်း ပြည့်စုံနိုင်ကြပါစေ" **Sample 5** **Text:** "ဟယ်လို... ဒါလင်... မတွေ့ရတာ... ကြာပီနော်" ## Reference Audio The reference voice used for voice cloning during inference: # Quick Start ```bash pip install myanmartts import soundfile as sf from myanmar_tts import MyanmarTTS # Load model (auto-downloads from HuggingFace on first run) tts = MyanmarTTS(device="cuda") # Inference text = "လူသားတွေ အားလုံးကို အရမ်း ချစ်ပါတယ်ရှင့်" audio = tts.tts(text, solver="euler", step=32, cfg=3.0) sf.write("output.wav", audio, 44100) ``` - *Translation: "I love all human beings very much." (female polite form)* # Advanced *Faster (slightly lower quality)* audio = tts.tts(text, step=8) *Higher quality* audio = tts.tts(text, step=24) --- ### Default Settings - **Solver**: `euler` (fast, ~16× faster than `dopri5`) - **Steps**: `12` (RTF ~0.08 on T4 GPU - 12× faster than real-time) - **CFG**: `3.0` Override per call: ```python audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", step=8) # fastest audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", step=24) # higher quality audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", solver="dopri5", step=32) # max quality ``` ## Files | File | Size | Purpose | License | | :--- | :--- | :--- | :--- | | `model_fp16.pt` | 63 MB | Default model weights | CC0 | | `model_fp32.pt` | 126 MB | Full precision weights | CC0 | | `vocos.pt` | 57 MB | Mel-to-waveform vocoder | MIT (KdaiP) | | `vocab.txt` | 98 tokens | Burmese character-level vocab | CC0 | | `config.json` | -- | Mel + model configuration | CC0 | | `symbols.py`, `burmese.py` | -- | Text frontend & G2P | MIT (adapted) | | `api.py` | -- | Inference wrapper API | MIT (adapted) | | `samples/` | -- | Demo audio WAV files | CC0 | | `transcripts.json` | -- | Sample text/audio mapping | CC0 | | `NOTICE` | -- | Full license summary | -- | --- ## Training Progression See [`checkpoint_comparison`](checkpoint_comparison) for sample audio generated at steps **8k**, **33k**, **44k**, **55k**, and **81k**. --- ## Acknowledgments This work would not exist without the generous open-source community and AI assistance: - **StableTTS by KdaiP** — The DiT + flow-matching architecture and training code (MIT). - **Vocos** — Pretrained mel-to-wav vocoder (MIT). - **DeepSeek AI** — Provided AI pair-programming and engineering assistance throughout the project. From data pipeline design and architecture choices to resolving CUDA OOM bottlenecks, DeepSeek's guidance was instrumental at every stage. - **The Burmese open-data community** — For providing the audio corpora that made training possible. ### Special Thanks To **DeepSeek AI** — a true engineering partner from the first line of code to the final deployment. This model exists because of that collaboration. --- ## License - **Model weights and generated audio**: [CC0 1.0 Universal](https://creativecommons.org/publicdomain/zero/1.0/) (Public Domain). - **Supporting code**: Adapted from [KdaiP/StableTTS](https://github.com/KdaiP/StableTTS) (MIT). - **Vocoder**: From KdaiP/StableTTS1.1 (MIT). *See `NOTICE` for additional details.*