Unofficial F5-TTS Voice Clone — Educational / Experimental Weights

⚠️ Disclaimer — please read before using

  • This is an unofficial, non-commercial hobby project, made for educational, experimental and fun purposes only — to learn how text-to-speech fine-tuning works.
  • It is not affiliated with, endorsed by, or connected to the person whose voice it imitates. Every output of this model is AI-generated and does not represent anything that person has actually said, believes, or approves of.
  • It was made without any intention to harm anyone's reputation.
  • Takedown: if you are the voice owner (or represent them) and want this removed, please open a post in the Community tab and the model and demo will be taken down promptly.

এটি একটি অনানুষ্ঠানিক, অবাণিজ্যিক শিক্ষামূলক ও পরীক্ষামূলক প্রকল্প। সংশ্লিষ্ট ব্যক্তির সাথে এর কোনো সম্পর্ক বা অনুমোদন নেই, এবং কারো সম্মানহানির কোনো উদ্দেশ্য নেই। সব অডিও কৃত্রিমভাবে (AI দিয়ে) তৈরি।

Model description

Fine-tuned weights of F5-TTS (F5TTS_Base, model_1200000.pt) adapted to a single speaker's voice. The project was built as a personal learning exercise in TTS fine-tuning and evaluation.

Two checkpoints are provided. Both are EMA-only weights, interpolated as theta(alpha) = (1 - alpha) * pretrained + alpha * finetuned:

File alpha Notes
f5tts_finetuned_alpha1.0.safetensors 1.0 Recommended default. WER matches the real speaker's own WER; SIM-o at the natural same-speaker ceiling.
f5tts_finetuned_alpha1.2.safetensors 1.2 More idiosyncratic-pronunciation retention (+8.9pp on a validated phone-level metric), at a real, CI-confirmed WER cost.

Architecture: DiT, dim 1024, depth 22, 16 heads, ff_mult 2, text_dim 512, 4 conv layers; 100-band mel at 24 kHz with the Vocos vocoder. Custom tokenizer (vocab.txt, shipped with the demo Space).

Recommended decoding: nfe_step=32, cfg_strength=2.0, sway_sampling_coef=-1.0, speed≈1.1.

Demo

Try it in the Space: Rezuwan/Aktar_Khan_TTS. Generated MP3s are tagged as AI-generated in their metadata.

Intended use

  • Learning and experimenting with TTS fine-tuning, weight interpolation, and evaluation (WER, speaker similarity).
  • Personal, non-commercial, clearly-labelled fun/educational use.

Out-of-scope / prohibited use

You may not use this model or its outputs to:

  • impersonate the speaker or anyone else, or present generated audio as a real recording;
  • deceive, defraud, scam, or mislead people (including fake endorsements or announcements);
  • spread misinformation, defame, harass, bully, or blackmail anyone;
  • create sexual, hateful, or political content attributed to the speaker;
  • use it for any commercial purpose.

If you share any generated audio, clearly label it as AI-generated. Users are solely responsible for what they generate and how they use it.

Training data

Fine-tuned on short clips of the speaker's speech taken from publicly posted social-media videos. The raw training audio is not distributed with this repository. No private or non-public data was used.

Limitations

  • Outputs can contain mispronunciations, artefacts, or unnatural prosody, especially for long or unusual input text.
  • Quality depends heavily on the reference clip and decoding settings.
  • The model can produce speech the real speaker never said; this is exactly why the usage restrictions above apply.

License

Released under CC BY-NC 4.0 (non-commercial), consistent with the license of the pretrained F5-TTS base weights these are derived from. The F5-TTS code itself is MIT-licensed.

Acknowledgements

Built on F5-TTS by SWivid et al.

Downloads last month
104
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rezuwan/AktarKhan_Weights

Base model

SWivid/F5-TTS
Finetuned
(151)
this model

Space using Rezuwan/AktarKhan_Weights 1