voxcpm-tn / README.md
Code-Quasar's picture
Upload README.md with huggingface_hub
64f7ff2 verified
|
Raw History Blame Contribute Delete
1.13 kB
metadata
language:
  - ar
license: apache-2.0
base_model: openbmb/VoxCPM2
pipeline_tag: text-to-speech
tags:
  - tunisian
  - derja
  - arabic-dialect
  - voxcpm
  - tts

Code-Quasar/voxcpm-tn

VoxCPM2 fine-tuned for Tunisian Derja.

  • base: openbmb/VoxCPM2 (2B, tokenizer-free, 48 kHz)
  • method: full fine-tuning, 147 steps
  • data: ~? h Tunisian read speech, 16 kHz mono

Important: the dialect tag

Every training transcript was prefixed with (Tunisian Dialect), so untagged text is out of distribution. Always prefix it:

from voxcpm import VoxCPM

model = VoxCPM.from_pretrained("Code-Quasar/voxcpm-tn", load_denoiser=False)
wav = model.generate(text="(Tunisian Dialect) عسلامة، شنوة أحوالك اليوم؟")

import soundfile as sf
sf.write("out.wav", wav, 48000)

On a GPU with under ~8 GB, disable compilation:

import os
os.environ["TORCHDYNAMO_DISABLE"] = "1"

Limitations

Trained on a small corpus of read speech, so expect limited prosodic range and weaker long-form phrasing. Derived from source corpora with their own licence terms; the voices belong to real speakers.