IndicF5 Marathi – Numbers v2
IndicF5-Marathi-Numbers-v2 is a fine-tuned version of ai4bharat/IndicF5 for Marathi text-to-speech, focused on reading numbers correctly: amounts (९८,००० रुपये), lakhs and crores (73 लाखांच्या), times (७:२०), dates and percentages. It does zero-shot voice cloning from a short reference clip, like the base model.
It was built for customer-facing Marathi speech in insurance and finance, such as premium quotes, policy details and callback times, where base models often skip or misread numbers.
| Developed by | Turtlemint |
| Model type | Non-autoregressive flow-matching TTS (F5-TTS Base, Diffusion Transformer) |
| Language | Marathi (mr) |
| Fine-tuned from | ai4bharat/IndicF5 |
| Checkpoint | model_48000.pt (48,000 updates) |
| Output | 24 kHz mono audio (Vocos vocoder) |
Highlights
- 🔢 Reads numbers in Marathi: trained on 5,000 extra sentences with Devanagari and Latin digits, comma-grouped amounts, times and dates in context.
- 🗣️ Zero-shot voice cloning: give a 5–10 s reference clip and its transcript.
- 🇮🇳 Natural Marathi prosody: keeps the base IndicF5 Marathi quality, further tuned on about 31.5 h of Marathi speech.
- ⚡ Fast, non-autoregressive inference: no autoregressive decoding, so speed stays stable for long sentences.
Quick start
Installation
pip install git+https://github.com/ai4bharat/IndicF5.git
pip install huggingface_hub soundfile
Inference
import soundfile as sf
from huggingface_hub import hf_hub_download
from f5_tts.model import DiT
from f5_tts.infer.utils_infer import (
load_model, load_checkpoint, load_vocoder,
preprocess_ref_audio_text, infer_process,
)
REPO_ID = "SwarajSolanke-turtle/Marathi_Text_To_Speech"
ckpt = hf_hub_download(REPO_ID, "model_48000.pt")
vocab = hf_hub_download(REPO_ID, "vocab.txt")
device = "cuda"
model_cfg = dict(dim=1024, depth=22, heads=16, ff_mult=2, text_dim=512, conv_layers=4)
model = load_model(DiT, model_cfg, mel_spec_type="vocos", vocab_file=vocab, device=device)
model = load_checkpoint(model, ckpt, device, use_ema=True)
vocoder = load_vocoder(vocoder_name="vocos", device=device)
ref_audio, ref_text = preprocess_ref_audio_text(
"reference.wav", # 5–10 s clean Marathi speech
"संदर्भ ऑडिओमध्ये बोललेला मजकूर येथे लिहा.", # its exact transcript
)
text = "७३ लाखांच्या कव्हरसाठी वार्षिक प्रीमियम सुमारे ९८,००० रुपये येतो."
wav, sr, _ = infer_process(ref_audio, ref_text, text, model, vocoder,
mel_spec_type="vocos", device=device)
sf.write("output.wav", wav, sr)
Tip: results are best when the reference clip is clean (no music or noise), 5–10 s long, and its transcript is exact.
Training details
Training data
| Dataset | Utterances | Duration | Description |
|---|---|---|---|
| Marathi speech (original) | 10,939 | 22.37 h | Clean Marathi read speech |
| Numbers v1 | 2,791 | 3.24 h | Synthetic sentences with numbers |
| Numbers v2 (new) | 5,000 | 5.95 h | Wider set of number formats: amounts, lakhs/crores, times, dates, mixed scripts |
| Total (merged) | 18,730 | 31.56 h |
Numbers v2 clips are 1.7–8.1 s long (mean 4.3 s). The sentences are templated, domain-specific Marathi text about insurance, payments and scheduling. Audio was generated with an earlier Marathi fine-tune of this model and mixed with real speech so the model doesn't drift.
Training procedure
| Hyperparameter | Value |
|---|---|
| Architecture | DiT: dim 1024, depth 22, 16 heads, FF mult 2, text dim 512, 4 conv layers |
| Objective | Conditional flow matching (masked mel infilling) |
| Tokenizer | Custom IndicF5 character vocabulary |
| Effective batch size | 32 samples (2 per GPU × 16 grad accumulation) |
| Optimizer | AdamW |
| Learning rate | 1e-5 with linear warmup and decay (≈9.65e-6 → 6.73e-6 over this stage) |
| Total updates | 48,000 |
| EMA | Yes (use EMA weights at inference) |
| Hardware | 1 × NVIDIA A10G (24 GB) |
| Training time (final stage) | ~2.8 h for updates 8k → 48k |
Evaluation
Training loss
Flow-matching loss on the training set, averaged over 4,000-update windows. Each step samples a random noise level, so single-step loss is noisy; the window averages are the meaningful numbers.
| Updates | Mean loss | Min loss |
|---|---|---|
| 8k – 12k | 0.5636 | 0.2397 |
| 12k – 16k | 0.5689 | 0.2238 |
| 16k – 20k | 0.5703 | 0.2393 |
| 20k – 24k | 0.5652 | 0.2313 |
| 24k – 28k | 0.5593 | 0.2317 |
| 28k – 32k | 0.5503 | 0.2186 |
| 32k – 36k | 0.5603 | 0.2247 |
| 36k – 40k | 0.5625 | 0.2267 |
| 40k – 44k | 0.5508 | 0.2236 |
| 44k – 48k | 0.5558 | 0.2171 |
| Summary at 48k updates | Value |
|---|---|
| Mean loss (last 1k updates) | 0.556 ± 0.377 |
| Best single-step loss | 0.2171 |
| Final learning rate | 6.73e-6 |
The loss levels off at about 0.55–0.56, so the model has converged at this checkpoint.
Objective metrics
Held-out intelligibility (CER/WER from a Marathi ASR model) and naturalness (MOS) results are not published yet. They will be added to this card when available.
Intended use
Intended for
- Marathi voice assistants, IVR systems and customer-communication audio
- Reading numbers aloud in Marathi: prices, premiums, policy numbers, dates, times
- Research on Indic TTS and text normalization
Out of scope
- Imitating real people without their explicit consent
- Fraud, impersonation, disinformation, or any audio presented as a real person's speech without disclosure
- Languages other than Marathi (the base model supports more Indic languages, but this fine-tune targets Marathi only)
Limitations and bias
- Synthetic number data: part of the training audio was produced by an earlier version of this model, so its artifacts and pronunciation habits may carry over.
- Domain bias: number sentences are mostly about insurance and finance, so general or literary text may sound less natural.
- Very long numbers: IDs, phone numbers and long digit strings are best expanded or spaced out before synthesis.
- Depends on the reference clip: a noisy or wrongly transcribed reference lowers quality and can cause skipped words.
- Code-mixed text: Marathi mixed with English words is only partly supported.
Ethical considerations
This model can clone voices from a few seconds of audio. Use it only with reference audio you have permission to use, and disclose clearly when audio is AI-generated. Do not use it to impersonate anyone or to deceive.
Citation
If you use this model, please cite the original F5-TTS and IndicF5 works:
@article{chen2024f5tts,
title = {F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching},
author = {Chen, Yushen and Niu, Zhikang and Ma, Ziyang and Deng, Keqi and Wang, Chunhui and Zhao, Jian and Yu, Kai and Chen, Xie},
journal = {arXiv preprint arXiv:2410.06885},
year = {2024}
}
@misc{AI4Bharat_IndicF5_2025,
author = {Praveen S V and Srija Anand and Soma Siddhartha and Mitesh M. Khapra},
title = {IndicF5: High-Quality Text-to-Speech for Indian Languages},
year = {2025},
url = {https://github.com/AI4Bharat/IndicF5}
}
Acknowledgements
- AI4Bharat for IndicF5
- SWivid/F5-TTS for the F5-TTS architecture and training code
- Vocos for the vocoder
Note
- model do not understand the complex number or the number on which model is not trained if this is case then performed the normalization of digit into words and then send to the model , it will then work well on the data you have
- Downloads last month
- 119
Model tree for SwarajSolanke-turtle/Marathi_Text_To_Speech
Base model
ai4bharat/IndicF5