voxcpm-tn / README.md
Code-Quasar's picture
Upload README.md with huggingface_hub
64f7ff2 verified
|
Raw History Blame Contribute Delete
1.13 kB
---
language: [ar]
license: apache-2.0
base_model: openbmb/VoxCPM2
pipeline_tag: text-to-speech
tags: [tunisian, derja, arabic-dialect, voxcpm, tts]
---
# Code-Quasar/voxcpm-tn
VoxCPM2 fine-tuned for **Tunisian Derja**.
- base: `openbmb/VoxCPM2` (2B, tokenizer-free, 48 kHz)
- method: full fine-tuning, 147 steps
- data: ~? h Tunisian read speech, 16 kHz mono
## Important: the dialect tag
Every training transcript was prefixed with `(Tunisian Dialect)`, so **untagged text is
out of distribution**. Always prefix it:
```python
from voxcpm import VoxCPM
model = VoxCPM.from_pretrained("Code-Quasar/voxcpm-tn", load_denoiser=False)
wav = model.generate(text="(Tunisian Dialect) عسلامة، شنوة أحوالك اليوم؟")
import soundfile as sf
sf.write("out.wav", wav, 48000)
```
On a GPU with under ~8 GB, disable compilation:
```python
import os
os.environ["TORCHDYNAMO_DISABLE"] = "1"
```
## Limitations
Trained on a small corpus of read speech, so expect limited prosodic range and
weaker long-form phrasing. Derived from source corpora with their own licence
terms; the voices belong to real speakers.