Spaces:
Running
Numerals as a stress case for Arabic TTS (open data)
Thanks for the arena. Sharing a small open measurement in case it's useful for prompt selection.
ArNum-TTS (CC BY 4.0) has 15 MSA sentences, each written three ways: 2026, ٢٠٢٦, and spelled out. With Whisper plus a number parser as the listener:
- fish s2.1-pro-free recovered 11/15 numbers in Western digits and 1/15 in Arabic-Indic digits.
- ArTST (MBZUAI/speecht5_tts_clartts_ar) recovered 0/30 digit numerals, because ٠–٩ aren't in its vocabulary.
It's only 45 utterances, scored automatically.
Would prompts with dates, prices and phone numbers in Arabic-Indic digits be a useful addition to the arena's prompt set?
@syamjithnk
We have a section in https://huggingface.co/spaces/Navid-AI/Arabic-TTS-Arena/blob/main/examples.py#L63 for edge cases regarding numbers/dates/abbreviations and they were interesting in seperating models.
will give your examples a look and may add couple of sentences from it to the arena official examples.
Thanks, Mohamed — I appreciate you taking a look. The dataset contains 15 matched MSA sentences in three forms (Western digits, Arabic-Indic digits and spelled-out numbers), under CC BY 4.0. Keeping the three variants together could help separate numeral-format effects from sentence content.
https://huggingface.co/datasets/syamjithnk/arnum-tts
One clarification to my earlier comment: the vocabulary gap applies specifically to Arabic-Indic digits; it does not by itself explain the failures on Western digits. The reported scores use Whisper plus a number parser, without a human listening pass, so I treat them as a small diagnostic rather than a general ranking. Thanks for considering the examples for the arena.