|
Download README.md from freococo/MyanmarTTS: direct link, hf CLI and curl.
- Browser
- Download file 6.59 kB
-
https://huggingface.co/freococo/MyanmarTTS/resolve/main/README.md
- Command line
-
hf download hf://freococo/MyanmarTTS/README.md
-
curl -L -o README.md https://huggingface.co/freococo/MyanmarTTS/resolve/main/README.md
6.59 kB
| license: cc0-1.0 | |
| language: | |
| - my | |
| tags: | |
| - text-to-speech | |
| - tts | |
| - burmese | |
| - myanmar | |
| - stabletts | |
| - from-scratch | |
| pipeline_tag: text-to-speech | |
| # MyanmarTTS | |
| A from-scratch Burmese (Myanmar) text-to-speech model. 31M parameters, ~63 MB (fp16). | |
| - **Language**: Burmese (မြန်မာဘာသာ) | |
| - **Training data**: ~1.34M samples (news + real-world audio) | |
| - **Training steps**: 81,000 | |
| - **Training hardware**: A100-80GB (Colab Pro+), ~13 hours, ~88 compute units | |
| - **Inference hardware**: Free Colab T4, any modern GPU, or CPU (slower) | |
| - **Architecture**: StableTTS (DiT + flow matching) + Vocos vocoder | |
| - **License**: **CC0 1.0** (public domain, no attribution required) | |
| --- | |
| ## Sample Output | |
| All samples generated with default settings (`euler` solver, 12 steps). | |
| **Sample 0** | |
| **Text:** "မြန်မာလူမျိုးများဟာ အလွန် ယဥ်ကျေးသိမ်မွေ့ပြီး ဧည့်သည်များကို ပျူပျူငှာငှာ လှိုက်လှိုက်လှဲလှဲနဲ့ ကြိုဆိုကြပါတယ်" | |
| <audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_0.wav"></audio> | |
| **Sample 1** | |
| **Text:** "မင်္ဂလာပါရှင် ကျွန်မကတော့ မြန်မာလူမျိုး ကရင်တိုင်းရင်းသူ အမျိုးသမီးလေး တစ်ဦး ဖြစ်ပါတယ်" | |
| <audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_1.wav"></audio> | |
| **Sample 2** | |
| **Text:** "ဒီနေ့ ကျွန်မတို့ရဲ့ တီတီအက်စ် စနစ်သစ်လေး မော်ဒယ်အသစ်လေးတစ်ခုကို အောင်မြင်စွာ လေ့ကျင့် သင်ကြားနိုင်ခဲ့ပါတယ်" | |
| <audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_2.wav"></audio> | |
| **Sample 3** | |
| **Text:** "လူသားတိုင်း လူသားတိုင်း ကိုယ်စိတ်နှစ်ဖြာ ကျန်းမာရွှင်လန်းပြီး စီးပွားလာဘ်လာဘတွေ ဒီရေအလား ကြီးပွား တိုးတက်နိုင်ကြပါစေ" | |
| <audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_3.wav"></audio> | |
| **Sample 4** | |
| **Text:** "ဒီအသံထုတ်စနစ်လေးကို အသုံးပြုသူတိုင်း ကျန်းမာချမ်းသာပြီး လိုရာဆန္ဒတွေ တလုံးတဝတည်း ပြည့်စုံနိုင်ကြပါစေ" | |
| <audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_4.wav"></audio> | |
| **Sample 5** | |
| **Text:** "ဟယ်လို... ဒါလင်... မတွေ့ရတာ... ကြာပီနော်" | |
| <audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_5.wav"></audio> | |
| ## Reference Audio | |
| The reference voice used for voice cloning during inference: | |
| <audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_0.wav"></audio> | |
| # Quick Start | |
| ```bash | |
| pip install myanmartts | |
| import soundfile as sf | |
| from myanmar_tts import MyanmarTTS | |
| # Load model (auto-downloads from HuggingFace on first run) | |
| tts = MyanmarTTS(device="cuda") | |
| # Inference | |
| text = "လူသားတွေ အားလုံးကို အရမ်း ချစ်ပါတယ်ရှင့်" | |
| audio = tts.tts(text, solver="euler", step=32, cfg=3.0) | |
| sf.write("output.wav", audio, 44100) | |
| ``` | |
| - *Translation: "I love all human beings very much." (female polite form)* | |
| # Advanced | |
| *Faster (slightly lower quality)* | |
| audio = tts.tts(text, step=8) | |
| *Higher quality* | |
| audio = tts.tts(text, step=24) | |
| --- | |
| ### Default Settings | |
| - **Solver**: `euler` (fast, ~16× faster than `dopri5`) | |
| - **Steps**: `12` (RTF ~0.08 on T4 GPU - 12× faster than real-time) | |
| - **CFG**: `3.0` | |
| Override per call: | |
| ```python | |
| audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", step=8) # fastest | |
| audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", step=24) # higher quality | |
| audio = tts.tts("ဟယ်လို ဒါလင် မတွေ့ရတာ ကြာပီ", solver="dopri5", step=32) # max quality | |
| ``` | |
| ## Files | |
| | File | Size | Purpose | License | | |
| | :--- | :--- | :--- | :--- | | |
| | `model_fp16.pt` | 63 MB | Default model weights | CC0 | | |
| | `model_fp32.pt` | 126 MB | Full precision weights | CC0 | | |
| | `vocos.pt` | 57 MB | Mel-to-waveform vocoder | MIT (KdaiP) | | |
| | `vocab.txt` | 98 tokens | Burmese character-level vocab | CC0 | | |
| | `config.json` | -- | Mel + model configuration | CC0 | | |
| | `symbols.py`, `burmese.py` | -- | Text frontend & G2P | MIT (adapted) | | |
| | `api.py` | -- | Inference wrapper API | MIT (adapted) | | |
| | `samples/` | -- | Demo audio WAV files | CC0 | | |
| | `transcripts.json` | -- | Sample text/audio mapping | CC0 | | |
| | `NOTICE` | -- | Full license summary | -- | | |
| --- | |
| ## Training Progression | |
| See [`checkpoint_comparison`](checkpoint_comparison) for sample audio generated at steps **8k**, **33k**, **44k**, **55k**, and **81k**. | |
| --- | |
| ## Acknowledgments | |
| This work would not exist without the generous open-source community and AI assistance: | |
| - **StableTTS by KdaiP** — The DiT + flow-matching architecture and training code (MIT). | |
| - **Vocos** — Pretrained mel-to-wav vocoder (MIT). | |
| - **DeepSeek AI** — Provided AI pair-programming and engineering assistance throughout the project. From data pipeline design and architecture choices to resolving CUDA OOM bottlenecks, DeepSeek's guidance was instrumental at every stage. | |
| - **The Burmese open-data community** — For providing the audio corpora that made training possible. | |
| ### Special Thanks | |
| To **DeepSeek AI** — a true engineering partner from the first line of code to the final deployment. This model exists because of that collaboration. | |
| --- | |
| ## License | |
| - **Model weights and generated audio**: [CC0 1.0 Universal](https://creativecommons.org/publicdomain/zero/1.0/) (Public Domain). | |
| - **Supporting code**: Adapted from [KdaiP/StableTTS](https://github.com/KdaiP/StableTTS) (MIT). | |
| - **Vocoder**: From KdaiP/StableTTS1.1 (MIT). | |
| *See `NOTICE` for additional details.* |