IndicF5 Marathi – Numbers v2

IndicF5-Marathi-Numbers-v2 is a fine-tuned version of ai4bharat/IndicF5 for Marathi text-to-speech, focused on reading numbers correctly: amounts (९८,००० रुपये), lakhs and crores (73 लाखांच्या), times (७:२०), dates and percentages. It does zero-shot voice cloning from a short reference clip, like the base model.

It was built for customer-facing Marathi speech in insurance and finance, such as premium quotes, policy details and callback times, where base models often skip or misread numbers.

Developed by Turtlemint
Model type Non-autoregressive flow-matching TTS (F5-TTS Base, Diffusion Transformer)
Language Marathi (mr)
Fine-tuned from ai4bharat/IndicF5
Checkpoint model_48000.pt (48,000 updates)
Output 24 kHz mono audio (Vocos vocoder)

Highlights

  • 🔢 Reads numbers in Marathi: trained on 5,000 extra sentences with Devanagari and Latin digits, comma-grouped amounts, times and dates in context.
  • 🗣️ Zero-shot voice cloning: give a 5–10 s reference clip and its transcript.
  • 🇮🇳 Natural Marathi prosody: keeps the base IndicF5 Marathi quality, further tuned on about 31.5 h of Marathi speech.
  • ⚡ Fast, non-autoregressive inference: no autoregressive decoding, so speed stays stable for long sentences.

Quick start

Installation

pip install git+https://github.com/ai4bharat/IndicF5.git
pip install huggingface_hub soundfile

Inference

import soundfile as sf
from huggingface_hub import hf_hub_download
from f5_tts.model import DiT
from f5_tts.infer.utils_infer import (
    load_model, load_checkpoint, load_vocoder,
    preprocess_ref_audio_text, infer_process,
)

REPO_ID = "SwarajSolanke-turtle/Marathi_Text_To_Speech"
ckpt  = hf_hub_download(REPO_ID, "model_48000.pt")
vocab = hf_hub_download(REPO_ID, "vocab.txt")

device = "cuda"
model_cfg = dict(dim=1024, depth=22, heads=16, ff_mult=2, text_dim=512, conv_layers=4)

model = load_model(DiT, model_cfg, mel_spec_type="vocos", vocab_file=vocab, device=device)
model = load_checkpoint(model, ckpt, device, use_ema=True)
vocoder = load_vocoder(vocoder_name="vocos", device=device)

ref_audio, ref_text = preprocess_ref_audio_text(
    "reference.wav",                       # 5–10 s clean Marathi speech
    "संदर्भ ऑडिओमध्ये बोललेला मजकूर येथे लिहा.",  # its exact transcript
)

text = "७३ लाखांच्या कव्हरसाठी वार्षिक प्रीमियम सुमारे ९८,००० रुपये येतो."
wav, sr, _ = infer_process(ref_audio, ref_text, text, model, vocoder,
                           mel_spec_type="vocos", device=device)

sf.write("output.wav", wav, sr)

Tip: results are best when the reference clip is clean (no music or noise), 5–10 s long, and its transcript is exact.


Training details

Training data

Dataset Utterances Duration Description
Marathi speech (original) 10,939 22.37 h Clean Marathi read speech
Numbers v1 2,791 3.24 h Synthetic sentences with numbers
Numbers v2 (new) 5,000 5.95 h Wider set of number formats: amounts, lakhs/crores, times, dates, mixed scripts
Total (merged) 18,730 31.56 h

Numbers v2 clips are 1.7–8.1 s long (mean 4.3 s). The sentences are templated, domain-specific Marathi text about insurance, payments and scheduling. Audio was generated with an earlier Marathi fine-tune of this model and mixed with real speech so the model doesn't drift.

Training procedure

Hyperparameter Value
Architecture DiT: dim 1024, depth 22, 16 heads, FF mult 2, text dim 512, 4 conv layers
Objective Conditional flow matching (masked mel infilling)
Tokenizer Custom IndicF5 character vocabulary
Effective batch size 32 samples (2 per GPU × 16 grad accumulation)
Optimizer AdamW
Learning rate 1e-5 with linear warmup and decay (≈9.65e-6 → 6.73e-6 over this stage)
Total updates 48,000
EMA Yes (use EMA weights at inference)
Hardware 1 × NVIDIA A10G (24 GB)
Training time (final stage) ~2.8 h for updates 8k → 48k

Evaluation

Training loss

Flow-matching loss on the training set, averaged over 4,000-update windows. Each step samples a random noise level, so single-step loss is noisy; the window averages are the meaningful numbers.

Updates Mean loss Min loss
8k – 12k 0.5636 0.2397
12k – 16k 0.5689 0.2238
16k – 20k 0.5703 0.2393
20k – 24k 0.5652 0.2313
24k – 28k 0.5593 0.2317
28k – 32k 0.5503 0.2186
32k – 36k 0.5603 0.2247
36k – 40k 0.5625 0.2267
40k – 44k 0.5508 0.2236
44k – 48k 0.5558 0.2171
Summary at 48k updates Value
Mean loss (last 1k updates) 0.556 ± 0.377
Best single-step loss 0.2171
Final learning rate 6.73e-6

The loss levels off at about 0.55–0.56, so the model has converged at this checkpoint.

Objective metrics

Held-out intelligibility (CER/WER from a Marathi ASR model) and naturalness (MOS) results are not published yet. They will be added to this card when available.


Intended use

Intended for

  • Marathi voice assistants, IVR systems and customer-communication audio
  • Reading numbers aloud in Marathi: prices, premiums, policy numbers, dates, times
  • Research on Indic TTS and text normalization

Out of scope

  • Imitating real people without their explicit consent
  • Fraud, impersonation, disinformation, or any audio presented as a real person's speech without disclosure
  • Languages other than Marathi (the base model supports more Indic languages, but this fine-tune targets Marathi only)

Limitations and bias

  • Synthetic number data: part of the training audio was produced by an earlier version of this model, so its artifacts and pronunciation habits may carry over.
  • Domain bias: number sentences are mostly about insurance and finance, so general or literary text may sound less natural.
  • Very long numbers: IDs, phone numbers and long digit strings are best expanded or spaced out before synthesis.
  • Depends on the reference clip: a noisy or wrongly transcribed reference lowers quality and can cause skipped words.
  • Code-mixed text: Marathi mixed with English words is only partly supported.

Ethical considerations

This model can clone voices from a few seconds of audio. Use it only with reference audio you have permission to use, and disclose clearly when audio is AI-generated. Do not use it to impersonate anyone or to deceive.


Citation

If you use this model, please cite the original F5-TTS and IndicF5 works:

@article{chen2024f5tts,
  title   = {F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching},
  author  = {Chen, Yushen and Niu, Zhikang and Ma, Ziyang and Deng, Keqi and Wang, Chunhui and Zhao, Jian and Yu, Kai and Chen, Xie},
  journal = {arXiv preprint arXiv:2410.06885},
  year    = {2024}
}

@misc{AI4Bharat_IndicF5_2025,
  author       = {Praveen S V and Srija Anand and Soma Siddhartha and Mitesh M. Khapra},
  title        = {IndicF5: High-Quality Text-to-Speech for Indian Languages},
  year         = {2025},
  url          = {https://github.com/AI4Bharat/IndicF5}
}

Acknowledgements

Note

  • model do not understand the complex number or the number on which model is not trained if this is case then performed the normalization of digit into words and then send to the model , it will then work well on the data you have
Downloads last month
119
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SwarajSolanke-turtle/Marathi_Text_To_Speech

Finetuned
(18)
this model

Paper for SwarajSolanke-turtle/Marathi_Text_To_Speech