Instructions to use NewGame/libritts_f5tts_training with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ESPnet
How to use NewGame/libritts_f5tts_training with ESPnet:
unknown model type (must be text-to-speech or automatic-speech-recognition)
- Notebooks
- Google Colab
- Kaggle
ESPnet3 f5tts model
Packed model bundle generated from egs3/libritts/f5tts.
Model
- Repository:
NewGame/libritts_f5tts_training - Recipe:
egs3/libritts/f5tts - Corpus:
libritts - System:
f5tts - Creator:
thanapattrachu - Created:
2026-10-08T15:00:52 - Branch:
espnet3/recipe/f5tts_libritts - Git:
481290130d6(clean) - Origin: https://github.com/NewGamezzz/espnet.git
Model summary
- Class:
F5TTS - Total parameters:
157,945,188 - Learnable parameters:
157,945,188(100.0%) - Non-trainable parameters:
0 - Parameter size:
631.78 MB - Buffers:
4,194,338(16.78 MB) - Modules:
418total,275leaf - DType composition:
torch.float32(102.7%), torch.int64(0.0%)
Usage
from espnet3.api.inference import load
model = load("NewGame/libritts_f5tts_training")
output = model(
"The text to synthesize.",
"reference.wav", # the voice to clone: a path, or a (rate, samples) pair
"The transcript of the reference recording.",
)
samples, sample_rate = output["wav"].array, output["wav"].rate
Packaging
- Bundle:
model_pack - Exp dir:
./exp/training - Strategy:
copy experiment outputs; include extra recipe assets; apply exclude filters
Results
| dataset | fwhisper_cer | fwhisper_cer_delete | fwhisper_cer_equal | fwhisper_cer_insert | fwhisper_cer_replace | fwhisper_wer | fwhisper_wer_delete | fwhisper_wer_equal | fwhisper_wer_insert | fwhisper_wer_replace | spk_similarity | utmos |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| librispeech_pc | 0.7644 | 0.228 | 94.4667 | 0.354 | 0.1429 | 1.6404 | 0.0293 | 17.9494 | 0.0728 | 0.1961 | 0.7058 | 4.1267 |
Training config
expand
num_device: 4
num_nodes: 1
task: null
recipe_dir: .
data_dir: ./data
exp_tag: training
exp_dir: ./exp/training
stats_dir: ./exp/training/stats
inference_dir: ./exp/training/inference
create_dataset:
recipe_dir: .
dataset:
_target_: espnet3.components.data.data_organizer.DataOrganizer
_recursive_: false
recipe_dir: .
train:
- data_src_args:
split: train
manifest_path: ./data/manifests_filtered/train.tsv
fs: 24000
valid:
- data_src_args:
split: valid
manifest_path: ./data/manifests_filtered/valid.tsv
fs: 24000
test: null
preprocessor:
_target_: espnet2.train.preprocessor.CommonPreprocessor
token_type: char
token_list: ./data/tokens/char_tokens.txt
text_cleaner: tacotron
g2p_type: null
_convert_: all
_convert_: all
model:
_target_: espnet3.systems.f5tts.f5tts.F5TTS
token_list: ./data/tokens/char_tokens.txt
feats_extract_config:
fs: 24000
n_fft: 1024
hop_length: 256
win_length: 1024
n_mels: 100
hidden_size: 768
depth: 18
attention_heads: 12
attention_head_size: 64
feed_forward_multiplier: 2
text_embedding_size: 512
text_mask_padding: false
rotary_attention_heads: 1
convolution_layers: 4
sigma: 0.0
audio_drop_probability: 0.3
condition_drop_probability: 0.2
mask_fraction_range:
- 0.7
- 1.0
ode_solver_method: euler
_convert_: all
optimizer:
_target_: torch.optim.AdamW
lr: 7.5e-05
betas:
- 0.9
- 0.999
weight_decay: 0.01
_convert_: all
scheduler:
_target_: espnet3.components.schedulers.linear_warmup_decay.LinearWarmupDecayLR
warmup_steps: 20000
total_steps: 600000
_convert_: all
scheduler_interval: step
scheduler_monitor: null
optimizers: null
schedulers: null
best_model_criterion:
- - valid/loss
- 1
- min
seed: null
init: null
parallel:
env: local
n_workers: 1
allow_overwrite_lock: false
dataloader:
collate_fn:
_target_: espnet2.train.collate_fn.CommonCollateFn
int_pad_value: 0
float_pad_value: 0.0
_convert_: all
train:
num_shards: 1
num_workers: 0
iter_factory:
_target_: espnet2.iterators.sequence_iter_factory.SequenceIterFactory
shuffle: true
collate_fn:
_target_: espnet2.train.collate_fn.CommonCollateFn
int_pad_value: 0
float_pad_value: 0.0
_convert_: all
batches:
type: numel
batch_size: 1
min_batch_size: 4
shape_files:
- ./exp/training/stats/train/feats_shape
batch_bins: 960000
_convert_: all
valid:
num_shards: 1
num_workers: 0
iter_factory:
_target_: espnet2.iterators.sequence_iter_factory.SequenceIterFactory
shuffle: false
collate_fn:
_target_: espnet2.train.collate_fn.CommonCollateFn
int_pad_value: 0
float_pad_value: 0.0
_convert_: all
batches:
type: numel
batch_size: 1
min_batch_size: 4
shape_files:
- ./exp/training/stats/valid/feats_shape
batch_bins: 960000
_convert_: all
trainer:
accelerator: auto
devices: 4
num_nodes: 1
strategy: auto
accumulate_grad_batches: 8
check_val_every_n_epoch: 1
gradient_clip_val: 1.0
log_every_n_steps: 50
logger:
- _target_: lightning.pytorch.loggers.TensorBoardLogger
save_dir: ./exp/training/tensorboard
name: tb_logger
_convert_: all
max_steps: 600000
max_epochs: 1000
callbacks:
- _target_: espnet3.components.callbacks.ema.EMACallback
decay: 0.9999
_convert_: all
fit: {}
sample_rate: 24000
n_mel_channels: 100
hop_length: 256
win_length: 1024
n_fft: 1024
remove_long_short:
min_wav_duration: 1.0
max_wav_duration: 20.0
splits:
- train
- valid
manifest_paths:
train: ./data/manifest/train.tsv
valid: ./data/manifest/valid.tsv
save_path: ./data/manifests_filtered
create_token_list:
manifest_path: ./data/manifests_filtered/train.tsv
save_path: ./data/tokens
filename: char_tokens.txt
token_type: char
cleaner: tacotron
g2p: null
non_linguistic_symbols: null
bpemodel: null
space_symbol: <space>
add_symbol:
- <blank>:0
- <unk>:1
- <sos/eos>:-1
vocabulary_size: -1
cutoff: 0
token_list: ./data/tokens/char_tokens.txt
Citing F5-TTS
@inproceedings{chen-etal-2025-f5,
author={Yushen Chen and Zhikang Niu and Ziyang Ma and Keqi Deng and Chunhui Wang
and Jian Zhao and Kai Yu and Xie Chen},
title={{F5-TTS}: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching},
year={2025},
booktitle={Proceedings of the 63rd Annual Meeting of the Association for
Computational Linguistics (Volume 1: Long Papers)},
pages={6255--6271},
doi={10.18653/v1/2025.acl-long.313}
}
Citing ESPnet
@inproceedings{watanabe2018espnet,
author={Shinji Watanabe and Takaaki Hori and Shigeki Karita and Tomoki Hayashi and
Jiro Nishitoba and Yuya Unno and Nelson Yalta and Jahn Heymann and Matthew Wiesner
and Nanxin Chen and Adithya Renduchintala and Tsubasa Ochiai},
title={{ESPnet}: End-to-End Speech Processing Toolkit},
year={2018},
booktitle={Proceedings of Interspeech},
pages={2207--2211},
doi={10.21437/Interspeech.2018-1456}
}
- Downloads last month
- 11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support