Unofficial F5-TTS Voice Clone — Educational / Experimental Weights
⚠️ Disclaimer — please read before using
- This is an unofficial, non-commercial hobby project, made for educational, experimental and fun purposes only — to learn how text-to-speech fine-tuning works.
- It is not affiliated with, endorsed by, or connected to the person whose voice it imitates. Every output of this model is AI-generated and does not represent anything that person has actually said, believes, or approves of.
- It was made without any intention to harm anyone's reputation.
- Takedown: if you are the voice owner (or represent them) and want this removed, please open a post in the Community tab and the model and demo will be taken down promptly.
এটি একটি অনানুষ্ঠানিক, অবাণিজ্যিক শিক্ষামূলক ও পরীক্ষামূলক প্রকল্প। সংশ্লিষ্ট ব্যক্তির সাথে এর কোনো সম্পর্ক বা অনুমোদন নেই, এবং কারো সম্মানহানির কোনো উদ্দেশ্য নেই। সব অডিও কৃত্রিমভাবে (AI দিয়ে) তৈরি।
Model description
Fine-tuned weights of F5-TTS (F5TTS_Base, model_1200000.pt) adapted to a single speaker's voice. The project was built as a personal learning exercise in TTS fine-tuning and evaluation.
Two checkpoints are provided. Both are EMA-only weights, interpolated as
theta(alpha) = (1 - alpha) * pretrained + alpha * finetuned:
| File | alpha | Notes |
|---|---|---|
f5tts_finetuned_alpha1.0.safetensors |
1.0 | Recommended default. WER matches the real speaker's own WER; SIM-o at the natural same-speaker ceiling. |
f5tts_finetuned_alpha1.2.safetensors |
1.2 | More idiosyncratic-pronunciation retention (+8.9pp on a validated phone-level metric), at a real, CI-confirmed WER cost. |
Architecture: DiT, dim 1024, depth 22, 16 heads, ff_mult 2, text_dim 512, 4 conv layers; 100-band mel at 24 kHz with the Vocos vocoder. Custom tokenizer (vocab.txt, shipped with the demo Space).
Recommended decoding: nfe_step=32, cfg_strength=2.0, sway_sampling_coef=-1.0, speed≈1.1.
Demo
Try it in the Space: Rezuwan/Aktar_Khan_TTS. Generated MP3s are tagged as AI-generated in their metadata.
Intended use
- Learning and experimenting with TTS fine-tuning, weight interpolation, and evaluation (WER, speaker similarity).
- Personal, non-commercial, clearly-labelled fun/educational use.
Out-of-scope / prohibited use
You may not use this model or its outputs to:
- impersonate the speaker or anyone else, or present generated audio as a real recording;
- deceive, defraud, scam, or mislead people (including fake endorsements or announcements);
- spread misinformation, defame, harass, bully, or blackmail anyone;
- create sexual, hateful, or political content attributed to the speaker;
- use it for any commercial purpose.
If you share any generated audio, clearly label it as AI-generated. Users are solely responsible for what they generate and how they use it.
Training data
Fine-tuned on short clips of the speaker's speech taken from publicly posted social-media videos. The raw training audio is not distributed with this repository. No private or non-public data was used.
Limitations
- Outputs can contain mispronunciations, artefacts, or unnatural prosody, especially for long or unusual input text.
- Quality depends heavily on the reference clip and decoding settings.
- The model can produce speech the real speaker never said; this is exactly why the usage restrictions above apply.
License
Released under CC BY-NC 4.0 (non-commercial), consistent with the license of the pretrained F5-TTS base weights these are derived from. The F5-TTS code itself is MIT-licensed.
Acknowledgements
Built on F5-TTS by SWivid et al.
- Downloads last month
- 104
Model tree for Rezuwan/AktarKhan_Weights
Base model
SWivid/F5-TTS