Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ViuAI
/
ViuAI_TTS_200M
Like
0
Text-to-Speech
PyTorch
Hindi
English
tts
flow-matching
diffusion-transformer
voice-cloning
zero-shot
elevenlabs-style
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
ViuAI_TTS_200M
33.1 GB
Ctrl+K
Ctrl+K
1 contributor
History:
94 commits
ViuAI
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
d62a3fe
verified
7 days ago
checkpoints
Upload checkpoints/viuai_tts_latest.pt with huggingface_hub
7 days ago
config
Update: training fixes, LICENSE + requirements + scripts
9 days ago
dataset
Turbo Speed: Vectorized FFT F0 pitch, cuDNN autotune, Tensor Core precision, speaker RAM cache & prefetch DataLoader
8 days ago
eval_benchmarks
Upload eval_benchmarks/epoch_60/02_hindi_female.wav with huggingface_hub
7 days ago
models
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
7 days ago
scripts
Master Fix: Real F0 pitch, acoustic Mel-80 to 100 filterbank vocoder projection, soundfile multi-format loader, safe weights_only
8 days ago
.gitattributes
Safe
4.92 kB
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
7 days ago
LICENSE
Safe
747 Bytes
Update: training fixes, LICENSE + requirements + scripts
9 days ago
README.md
Safe
4.09 kB
Initial commit: ViuAI_TTS_200M architecture, prosody engine, and weights
10 days ago
ROADMAP.md
Safe
2.32 kB
Initial commit: ViuAI_TTS_200M architecture, prosody engine, and weights
10 days ago
inference.py
Safe
9.18 kB
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
7 days ago
requirements.txt
Safe
185 Bytes
Performance: 32 parallel download workers and hf_transfer for 10x faster dataset downloading
8 days ago
test_viuai_tts_kaggle.ipynb
Safe
6.89 kB
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
7 days ago
train.py
Safe
45.9 kB
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
7 days ago
train_colab.ipynb
Safe
3.65 kB
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
7 days ago
training_epoch60_completed_memory.png
Safe
279 kB
xet
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
7 days ago
upload_to_hf.py
Safe
2.36 kB
Fix: Peak audio normalization (-1 dB), robust speaker search, and tuned CFG for crystal-clear benchmark generation
7 days ago
verify_model.py
Safe
4.47 kB
Master Fix: Real F0 pitch, acoustic Mel-80 to 100 filterbank vocoder projection, soundfile multi-format loader, safe weights_only
8 days ago